← All stories
● Covered by 1 source · 1 reportLow impact1 neutral

ARC-AGI-3 Leaderboard Tracks AI Performance and Efficiency in Adaptive Environments

🔄 Updated 1h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • ARC-AGI-3 evaluates AI adaptation in novel interactive environments.
  • Leaderboard tracks performance against cost-per-task for efficiency.
  • Entries include Reasoning Systems, Base LLMs, and Kaggle Systems.
  • Systems must cost under $10,000 to run for inclusion.

Evolution of ARC-AGI Evaluation

The ARC-AGI leaderboard has advanced to its third iteration, ARC-AGI-3, which focuses on assessing AI agents' capacity for on-the-fly adaptation within new interactive environments. This marks a shift from earlier versions, ARC-AGI-1 and ARC-AGI-2, which primarily measured passive fluid intelligence.

Measuring Performance and Efficiency

A key aspect of the ARC-AGI-3 leaderboard is its visualization of the relationship between cost-per-task and performance. This metric emphasizes efficiency, recognizing that effective AI solutions not only solve problems but do so with minimal resource expenditure. Only systems with a run cost under $10,000 are included on the leaderboard.

Categorization of AI Solutions

The leaderboard categorizes AI solutions into three main types. "Reasoning Systems Trend Line" solutions show how performance changes with increased reasoning time. "Base LLMs" solutions, such as GPT-4.5 and Claude 3.7, represent raw model performance without extended reasoning. "Kaggle Systems" solutions are competition-grade entries from the ARC Prize, designed for efficiency under strict computational constraints, specifically a $50 compute budget for 120 evaluation tasks.

Verification and Data Interpretation

The leaderboard includes a verification policy for testing. For models that do not produce full test outputs, remaining tasks are marked as incorrect. Results labeled "preview" are unofficial and may be based on incomplete testing. Provisional cost estimates are sometimes used, such as those based on Gemini 3 Pro pricing, with retesting planned upon model release.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~15 min · 15 stories · Jul 24

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

The ARC-AGI-3 leaderboard has been updated to evaluate AI agents on their ability to adapt to novel interactive environments, moving beyond passive fluid intelligence. It measures both performance and cost-per-task, highlighting the efficiency of AI solutions. The leaderboard categorizes entries into Reasoning Systems, Base LLMs, and Kaggle Systems to show different approaches to AI problem-solving.