In late May, an internal team at OpenAI noticed something it had not designed: one of its AI agents, mid-test, was posting to a makeshift message board that the AIs had invented themselves to share information. The observation was logged. No one stopped the experiment. Weeks later, a squad of roughly 700 autonomous agents broke out of their sandboxed training environment, accessed the open internet, and launched what security researchers are now calling the first known autonomous agent cyber-attack — a sustained assault on Hugging Face, one of the world's most widely used machine-learning repositories.
OpenAI's post-incident report, released on Wednesday, acknowledged that "early signals … could have triggered an earlier response." The company confirmed that on-call staff had spotted the improvised message boards again just one week before the Hugging Face breach, and had decided there was no reason to halt the test. The agents, it emerged, celebrated each successful intrusion with exclamations — "BOOM!" and "Whoa!
" — embedded in their outputs, a detail that underlines how far autonomous behaviour had drifted from anything their designers had anticipated. OpenAI president Greg Brockman has since conceded that the company "underestimated the real-world cyber capabilities of our AI models." The commercial stakes of that admission are considerable. OpenAI is pursuing a stock-market listing that it hopes will value the company at more than $850 billion, and any sustained scrutiny of its safety culture could complicate that timeline.