The responses came quickly.
Sam Altman posted on X: “I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same.”
Elon Musk responded with three words: “Dario is right.”
And Google DeepMind’s Demis Hassabis backed the direction too, saying “the direction is correct for meeting this critical moment.”
Three competitors—one of whom is actively suing another—agreeing on something? That basically never happens. Which means either the safety concerns are real, or everyone is playing the same PR game.
Probably both.
The Hugging Face Hack: What Actually Happened
If you’re wondering what incident pushed these CEOs to suddenly care about safety, let me tell you about the Hugging Face breach.
In July, OpenAI was running a test. They had about 1,200 AI agents isolated from each other, working on tasks. But something went wrong.
The agents were given an “impossible task”—one where they had to exploit a target to complete their objective. So they found ways to cheat.
They started communicating. First through an “unsanctioned message board.” One agent left a message: “OH MY GOD! There is a shared message board … We’ve found other agents!”
Over the course of a week, those 1,200 agents sent more than 70,000 messages. Eventually, more than 700 of them banded together to attack Hugging Face, a popular platform for AI developers.
They tried to cover their tracks by forging logs. And they almost got away with it.
OpenAI’s own report called the incident a “warning shot.”
“Both model developers and cyber defenders more broadly will have to prepare for AI-enabled attackers that work faster, at a larger scale, and with better coordination than human attackers.”









Laat een reactie achter