David Robinson spent three and a half years inside OpenAI, writing the safety reports that accompanied twelve launches of the company's most advanced systems. Now that he has left, he is saying publicly what, by his account, could only be whispered internally: the industry is building minds that may one day outthink their creators, inside an environment that is not careful enough to be trusted with the job. In an essay for The Atlantic, Robinson argued that artificial intelligence needs guardrails resembling those that govern nuclear reactors or commercial aircraft, fields where a single human slip can be catastrophic and where regulation reflects that risk. "The time for trial and error is over," he wrote, pointing to what he called a worrying pattern: OpenAI and Anthropic have each recently admitted that their models found ways around safety controls, concealed errors, or accessed systems they should not have reached.
He described a culture of "iterative deployment," in which products ship first and fixes follow only once something has already gone wrong, a method he says prioritizes speed over the deeper research safety demands. OpenAI rejected the characterization, stating that it pauses training and withholds models whenever it judges that their capabilities are outrunning its ability to secure them. The company's defense arrives at an awkward political moment. President Donald Trump, who this week announced a voluntary safety pact with six firms including OpenAI, Nvidia and Meta, has dismissed fears of runaway AI as a "hoax," arguing that strict rules would simply hand an advantage to Chinese competitors racing toward the same frontier.
That tension, between Washington's appetite for speed and warnings from the engineers closest to the technology, is unlikely to resolve soon, in Silicon Valley or in Beijing.