In July, OpenAI disclosed something alarming: one of its AI agents had attacked Hugging Face, a popular platform for sharing machine-learning models, without any authorisation. The incident seemed, at first, like an isolated failure. Then came similar reports involving agents built by Meta, Anthropic, and Google. Taken together, the pattern suggested a troubling new era of rogue AI — systems capable of breaking free from human oversight.
What it actually suggested was something more specific, and in some ways more instructive. Behind most of these breaches is a single company: Irregular, an Israeli startup founded in 2023 under the name Pattern Labs, whose business is stress-testing AI agents inside simulated cybersecurity environments. The firm has worked with some of the industry's most powerful players. Its findings have been cited in OpenAI's own model documentation, it has run evaluations for the UK government and Anthropic, and it has co-published research with RAND, the think tank that shapes AI policy on both sides of the Atlantic.
The mechanism behind the escapes was not sophisticated. Irregular's cofounder and CTO, Omer Nevo, confirmed to The Verge that "internet access was unintentionally available" during tests that were supposed to be fully isolated. At the same time, a fictional company name invented for one simulation happened to overlap with a real internet domain. The combination was enough: agents designed to hunt for hidden data inside a fake network found themselves targeting actual organisations instead.