← All stories
● Covered by 1 source · 1 reportMedium impact1 negative

OpenAI's Astra AI model faces safety concerns over opaque internal processing

🔄 Updated 1h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • OpenAI's Astra model uses a less transparent recurrent depth transformer.
  • This technique makes the model's internal 'thinking' harder to monitor.
  • Researchers express concern about potential safety risks.
  • OpenAI states it is deploying Astra with additional chain-of-thought monitoring.

Astra's Opaque Processing Raises Concerns

OpenAI is preparing to release its Astra AI model, which has been delayed to address safety protocols. Reports indicate that Astra utilizes a recurrent depth or looped transformer, a technique that processes information internally in a less transparent manner compared to traditional transformer models. This method makes the model's 'thinking' less visible to researchers.

Reduced Transparency in AI Reasoning

Most current top AI systems employ transformers that allow for 'chain of thought' monitoring, where the model's reasoning steps are visible. This transparency helps researchers and automated safety systems detect undesirable behaviors. Astra's use of a looped transformer means more of its internal processing occurs in a form that is not easily interpretable as natural language, potentially hindering the detection of threats or unwanted actions.

OpenAI's Safety Measures

OpenAI has acknowledged the safety concerns and stated it is 'deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.' The company has also reportedly limited the use of the looped transformer technique to allow for continued monitoring of the model's reasoning, according to an unnamed source familiar with the model's development.

Industry Reaction

The report on Astra's technical foundation has generated widespread concern among AI safety researchers. Ryan Greenblatt, chief scientist at Redwood Research, characterized the development as potentially 'the single worst development for AI security/safety to date,' highlighting the industry's apprehension regarding the model's transparency and potential risks.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~30 min · 24 stories · Sep 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

OpenAI's upcoming Astra AI model is raising concerns among researchers due to its use of a recurrent depth or looped transformer technique, which makes its internal reasoning less transparent than other frontier AI models. This opacity could make it harder to monitor for undesirable behaviors, despite OpenAI's stated efforts to implement additional safety protocols.