← All stories
● Covered by 1 source · 1 reportMedium impact1 positive

Simular's Sai Agent Achieves 73% Success on OSWorld 2.0 Benchmark

🔄 Updated 2h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Sai agent scored 73% on OSWorld 2.0 benchmark.
  • Outperformed GPT-5.6 Sol (62.57%) and Opus 5 (70.57%).
  • Sai operates at approximately 2/3 the cost of competitors.
  • Designed for everyday workplace tasks, not just raw throughput.

Sai Agent Tops OSWorld 2.0 Benchmark

Simular's Sai agent achieved a 73% success rate on the OSWorld 2.0 benchmark, as announced in a recent update. This benchmark consists of 108 tasks designed to evaluate an agent's ability to perform lengthy professional tasks that typically take skilled humans over an hour to complete.

Performance Against Competitors

Sai's 73% score places it ahead of OpenAI's GPT-5.6 Sol, which scored 62.57%, and Anthropic's Opus 5, which scored 70.57%. Simular also stated that Sai operates at roughly two-thirds the cost of these competing models.

Focus on Real-World Workplace Tasks

Sai is built to handle real-world workplace functions, operating on full desktop applications and webpages, calling APIs, and writing code. Simular emphasizes that the agent is designed for routine, necessary work like recruitment outreach or invoice validation, rather than solely for achieving high throughput in academic benchmarks.

Simular's co-founder and CTO, Jiachen Yang, highlighted the importance of cost-effectiveness for agents performing everyday tasks, suggesting that high prices for underlying models are unnecessary for such applications. He believes Sai's cost-outcome tradeoff on OSWorld 2.0 is a step towards more accessible AI solutions.

Technology and Company Background

The Sai agent combines frontier and specialist models, using dedicated interfaces to perceive and act on a user's computer. Simular, founded by ex-DeepMind scientists Ang Li (CEO) and Jiachen Yang (CTO), describes itself as a "research-focused agentic startup." The company employs a neurosymbolic method, integrating neural networks' flexibility with symbolic code's precision.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~20 min · 17 stories · Aug 27

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Simular's Sai agent scored 73% on the OSWorld 2.0 benchmark, outperforming OpenAI's GPT-5.6 Sol and Anthropic's Opus 5. This result indicates progress in AI agents handling complex, real-world professional tasks at a lower operational cost.