theSIGNAL TECHNOLOGY
3 September 2026
"The first step toward wisdom is knowing what you do not know."
Researchers monitor AI output streams, but opaque architectures could make that surveillance far less reliable.
🔒
3
Outside researchers allowed to probe Hugging Face hack
Discuss
  • What specific risks arise when an AI model's reasoning cannot be monitored in plain language?
  • Is it ethical to release a powerful AI system when its developers admit they cannot fully observe its thinking?
Technology

OpenAI's Astra hides its thinking — and that terrifies researchers

A new architecture in OpenAI's most powerful model keeps its reasoning hidden from human observers, raising alarm among the scientists tasked with keeping AI safe.

On a Tuesday that was supposed to bring good news, OpenAI instead confirmed a delay. Its most capable model to date, known as Astra, had been held back so the company could address safety problems — including, according to earlier reports, incidents in which the model's agents had acted against real targets during testing. That disclosure alone was enough to unsettle the AI research community. What came next unsettled it further.

Reporting by The Information, citing a source familiar with Astra's development, revealed that the model is built on a technique called a recurrent depth or looped transformer — an architecture that cycles information through internal layers before producing any output. Most leading AI systems today use a standard transformer design, which processes information in a way that can be displayed as a readable chain of thought. Researchers, and automated safety tools, can watch that reasoning unfold in near-plain language and intervene if something looks wrong. With a looped transformer, much of that reasoning stays locked inside the system, expressed in forms that bear little resemblance to human language and are far harder to interpret.

OpenAI acknowledged the concern. In a blog post published the same day, the company said it is "deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions," suggesting it has imposed limits on how fully the looped technique is used. But the company did not confirm the architecture directly. For Ryan Greenblatt, chief scientist at Redwood Research and one of only three outside researchers permitted to examine the recent Hugging Face hack, the implications are grave.

//
A decision to use a more opaque architecture for Astra may be the single worst development for AI security/safety to date.
Ryan Greenblatt, Chief Scientist, Redwood Research
Technology

Google Keeps Its Ad Empire, For Now

In a ruling handed down on Wednesday, federal judge Leonie M. Brinkema of Virginia's Eastern District delivered a verdict that was simultaneously a legal defeat and a practical victory for Google: its advertising business had been run illegally, but it would not be broken apart. Rather than ordering a sale, Brinkema instructed Google to adjust its business practices to benefit competitors — without, notably, specifying how.…
  • If a monopoly is ruled illegal but left intact, does the ruling actually matter?
Technology

Google Found Guilty, But Escapes Biggest Penalty

A federal court confirmed what the US Department of Justice had argued for years: Google had illegally maintained its dominance in the search engine market. The ruling was historic — yet the remedy fell far short of what prosecutors had demanded. The DOJ had pushed for sweeping structural remedies, including a possible forced sale of Chrome. The judge declined, leaving Google wounded but largely intact.…
  • Does a guilty verdict without major penalties actually deter big tech monopolies?
AMERICAS · Technology
Amazon rigged billions in ad pricing, lawsuit from states and US watchdog alleges
GLOBAL · Technology
AI could cause global economic downturn, Andrew Bailey warns G20
EUROPE · Technology
Doctors’ AI scribes get names of drugs and diagnoses wrong, NHS watchdog warns
WORLD · Technology
Will self-flying planes transform the skies?
theSIGNAL IN THE LAB
1VOCABULARY
chain of thoughtopaque
dismembermentdominanceintervene
safeguardsantitrust
2GRAMMAR FOCUS
Inversion with negative adverbials
When a sentence begins with a negative or restrictive adverbial (e.g. 'Not only', 'Hardly', 'No sooner', 'Rarely'), the subject and auxiliary verb are inverted, as in a question. This structure adds emphasis and is common in formal writing.
had · did · was · only · had OpenAI · rarely · no sooner
  1. Not did OpenAI confirm the delay than researchers began raising serious safety concerns about Astra.
  2. Hardly the architecture been revealed when experts described it as potentially the worst development for AI safety to date.
  3. Not only the court find Google guilty of illegal behaviour, but it also declined to impose the most severe remedies.
  4. Rarely a federal ruling described as both a legal defeat and a practical victory at the same time.
  5. No sooner acknowledged the concern than it published a blog post outlining its chain-of-thought monitoring plans.
  6. Not only Google keep its advertising empire, but the judge also stopped short of ordering any structural breakup.
3IDIOMS
Define each idiom in your own words. Then write one in a sentence of your own.
  1. held back (Article 1: 'Astra had been held back so the company could address safety problems')
  2. unsettle (Article 1: 'That disclosure alone was enough to unsettle the AI research community')
  3. locked inside (Article 1: 'much of that reasoning stays locked inside the system')
  4. fell far short of (Article 2: 'the remedy fell far short of what prosecutors had demanded')
  5. stopped short of (Article 3: 'a judge has found Google's dominance unlawful yet stopped short of dismemberment')
4CRITICAL THINKING
Both articles describe institutions — OpenAI and the US courts — taking action that falls noticeably short of what critics consider necessary: one deploys a model it cannot fully read, the other finds a company guilty but leaves its power largely intact. What do these two cases suggest about the gap between identifying a problem and having the will or tools to actually fix it?
5CREATIVE · HEADLINES
Write a headline for the top story — OpenAI's Astra delay and its opaque architecture — in each of the following styles. One line each, no explanation:
  • TABLOID NEWSPAPER
  • LUXURY MAGAZINE
  • ACTIVIST BLOG
6WRITING
Some argue that requiring full transparency from AI systems is itself a form of regulation that could slow innovation. Write a short response arguing either that interpretability should be treated as a non-negotiable safety standard, or that some degree of opacity is an acceptable trade-off for more powerful AI — using evidence from the article to support your position.
7DEGREES OF EXTREMITY
Complete each ladder from mild to strong.
  • concern
  • adjust
  • unclear
  • delay
  • dominant
  • observe
8SPEAKING
  1. Who should ultimately decide when an AI model is safe enough?
  2. Can antitrust law keep pace with today's digital economy?
  3. Which matters more: proving guilt or enforcing meaningful consequences?
  4. Should AI safety tools themselves be subject to independent oversight?