On a Tuesday in June, a researcher named Jacob Coxon posted a message to social media that did not read like a publicity stunt. Coxon, who had just quit his position at Anthropic — the San Francisco-based AI company valued at roughly $61 billion — wrote that neither Anthropic nor his previous employer, OpenAI, was acting responsibly. Both, he argued, were "racing straight to self-improving superintelligence and gambling with our lives." What made the post unusual was not its tone, which is familiar in AI-safety circles, but its source: a person who had sat inside one of the labs doing the building.
The alarm was quickly amplified by two colleagues who were still on Anthropic's payroll. Evan Hubinger, who leads the company's alignment division — the team responsible for ensuring AI models behave in accordance with human goals — confirmed that Coxon was "correct." He went further, stating that he personally estimates a greater-than-10% probability that AI kills all humans within the next decade. That figure is striking not because it is a majority view, but because it comes from someone whose job it is to prevent exactly that outcome.
Samuel Marks, Anthropic's scalable oversight lead, added a separate analysis in a personal capacity, noting that concern tends to rise with seniority: the more a researcher knows, the more worried they appear to be. Anthropic, for its part, issued a statement defending its record on safety, citing "some of the strongest safeguards in the industry" and its role as a pioneer in interpretability research — the scientific effort to understand what AI models are actually computing inside. The company did not dispute the substance of its employees' fears.