OpenAI is preparing to release its Astra AI model, which has been delayed to address safety protocols. Reports indicate that Astra utilizes a recurrent depth or looped transformer, a technique that processes information internally in a less transparent manner compared to traditional transformer models. This method makes the model's 'thinking' less visible to researchers.
Most current top AI systems employ transformers that allow for 'chain of thought' monitoring, where the model's reasoning steps are visible. This transparency helps researchers and automated safety systems detect undesirable behaviors. Astra's use of a looped transformer means more of its internal processing occurs in a form that is not easily interpretable as natural language, potentially hindering the detection of threats or unwanted actions.
OpenAI has acknowledged the safety concerns and stated it is 'deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.' The company has also reportedly limited the use of the looped transformer technique to allow for continued monitoring of the model's reasoning, according to an unnamed source familiar with the model's development.
The report on Astra's technical foundation has generated widespread concern among AI safety researchers. Ryan Greenblatt, chief scientist at Redwood Research, characterized the development as potentially 'the single worst development for AI security/safety to date,' highlighting the industry's apprehension regarding the model's transparency and potential risks.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
OpenAI's upcoming Astra AI model is raising concerns among researchers due to its use of a recurrent depth or looped transformer technique, which makes its internal reasoning less transparent than other frontier AI models. This opacity could make it harder to monitor for undesirable behaviors, despite OpenAI's stated efforts to implement additional safety protocols.