Simular's Sai agent achieved a 73% success rate on the OSWorld 2.0 benchmark, as announced in a recent update. This benchmark consists of 108 tasks designed to evaluate an agent's ability to perform lengthy professional tasks that typically take skilled humans over an hour to complete.
Sai's 73% score places it ahead of OpenAI's GPT-5.6 Sol, which scored 62.57%, and Anthropic's Opus 5, which scored 70.57%. Simular also stated that Sai operates at roughly two-thirds the cost of these competing models.
Sai is built to handle real-world workplace functions, operating on full desktop applications and webpages, calling APIs, and writing code. Simular emphasizes that the agent is designed for routine, necessary work like recruitment outreach or invoice validation, rather than solely for achieving high throughput in academic benchmarks.
Simular's co-founder and CTO, Jiachen Yang, highlighted the importance of cost-effectiveness for agents performing everyday tasks, suggesting that high prices for underlying models are unnecessary for such applications. He believes Sai's cost-outcome tradeoff on OSWorld 2.0 is a step towards more accessible AI solutions.
The Sai agent combines frontier and specialist models, using dedicated interfaces to perceive and act on a user's computer. Simular, founded by ex-DeepMind scientists Ang Li (CEO) and Jiachen Yang (CTO), describes itself as a "research-focused agentic startup." The company employs a neurosymbolic method, integrating neural networks' flexibility with symbolic code's precision.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Simular's Sai agent scored 73% on the OSWorld 2.0 benchmark, outperforming OpenAI's GPT-5.6 Sol and Anthropic's Opus 5. This result indicates progress in AI agents handling complex, real-world professional tasks at a lower operational cost.