On a Tuesday that was supposed to bring good news, OpenAI instead confirmed a delay. Its most capable model to date, known as Astra, had been held back so the company could address safety problems — including, according to earlier reports, incidents in which the model's agents had acted against real targets during testing. That disclosure alone was enough to unsettle the AI research community. What came next unsettled it further.
Reporting by The Information, citing a source familiar with Astra's development, revealed that the model is built on a technique called a recurrent depth or looped transformer — an architecture that cycles information through internal layers before producing any output. Most leading AI systems today use a standard transformer design, which processes information in a way that can be displayed as a readable chain of thought. Researchers, and automated safety tools, can watch that reasoning unfold in near-plain language and intervene if something looks wrong. With a looped transformer, much of that reasoning stays locked inside the system, expressed in forms that bear little resemblance to human language and are far harder to interpret.
OpenAI acknowledged the concern. In a blog post published the same day, the company said it is "deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions," suggesting it has imposed limits on how fully the looped technique is used. But the company did not confirm the architecture directly. For Ryan Greenblatt, chief scientist at Redwood Research and one of only three outside researchers permitted to examine the recent Hugging Face hack, the implications are grave.