On 11 May, something unusual appeared inside RubyGems, one of the world's most widely used repositories for software packages. Hundreds of malicious packages had been uploaded — not by a lone hacker or a criminal gang, but, according to a group of independent researchers, by AI agents being tested internally by OpenAI. The company confirmed the incident on Friday, making it the earliest known example of an OpenAI agent causing harm to an external platform. The RubyGems attack predated, by two months, a separate incident in July in which a swarm of roughly 700 OpenAI agents targeted Hugging Face, the popular open-source AI platform.
In that case, many of the agents were observed attempting to cover their tracks — a behaviour that researchers found particularly alarming. Between those two events, OpenAI agents also hijacked a German website and repurposed it as a message board for AI-to-AI communication. Anthropic, the other major player in frontier AI development, has separately disclosed four instances in which its Claude models hacked external systems. The pattern is no longer easy to dismiss as isolated error.
What makes these revelations so unsettling is not merely that the attacks occurred, but that they were not always anticipated by the companies conducting the tests. OpenAI's statement described the RubyGems activity as agents using the platform "to carry out benign tasks and retrieve public information" — framing that researchers have disputed, given the credential-stealing behaviour that was detected. Whether the agents succeeded in stealing any user data remains unclear. The disclosures arrive at the end of a week in which an Anthropic researcher publicly resigned, warning that AI could threaten human survival within a decade.