OpenAI's new GPT-6 Astra model achieved a 99.9% score on the ARC-AGI-3 benchmark when evaluated with a provider adapter harness. This harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work. With a standard harness, Astra scored 62.7% on ARC-AGI-3 Semi-Private.
This performance represents a substantial increase compared to its predecessor, GPT-5.6 Sol, which scored 7.8% on the same benchmark. The ARC-AGI-3 benchmark is designed to test agentic intelligence in novel, abstract, turn-based environments, requiring models to explore, infer goals, and build internal models without explicit instructions.
GPT-6 Astra demonstrated the ability to turn unfamiliar environments into compact symbolic world models, representing game mechanics as logical rules and developing its own domain-specific language shorthand for tracking state and planning actions. The model also surpassed human action efficiency on ARC-AGI-3, using fewer actions than the median tested human on 96% of levels.
OpenAI describes GPT-6 Astra as a "generational leap" in capability, offering faster and more reliable computer use for advanced workflows. The company states that Astra is state-of-the-art in areas such as computer use, browsing, software engineering, cybersecurity, science, and professional work. It also saturates FrontierMath Tier 4 with a 98% score and ExploitBench with a 100% score.
The evaluation of Astra on ARC-AGI-3 involved specific settings within OpenAI's Responses API harness. These changes were made to reflect how the model performs in real-world use. The benchmark's purpose is to assess models' ability to navigate uncharted interactive territory and work out environment mechanics independently, rather than relying solely on training data.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
OpenAI has launched GPT-6 Astra, its latest AI model, which the company describes as a "generational leap" in capability, offering faster and more reliable computer use for advanced workflows. This release is significant as it claims state-of-the-art performance across various domains, potentially setting new benchmarks for AI model performance and application.
OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 Semi-Private benchmark using a standard harness, and 99.9% with a provider adapter harness. GPT-6 Astra surpassed human action efficiency on 96% of levels, demonstrating its ability to create symbolic world models and domain-specific language.
OpenAI's GPT-6 Astra model scored 98.6% on the ARC-AGI-3 benchmark, a significant increase from its predecessor GPT-5.6 Sol's 7.8%. This improvement demonstrates progress in AI models' ability to navigate unfamiliar interactive environments, though the evaluation setup for Astra differed from other models.