The leaderboard for AI models on Software Engineering (SWE) tasks has undergone multiple updates between December 2025 and July 2026. These updates include the addition of new models and the deprecation of others, indicating continuous evolution in the field of AI for software development.
Recent additions to the leaderboard include GLM 5.2, DeepSeek-V4 Pro, DeepSeek-V4 Flash, MiMo V2.5 Pro, Qwen3.6-35B-A3B, Qwen3.6-27B, Gemma 4 31B, Gemini 3.5 Flash, MiniMax M3, Claude Opus 4.8, gpt-5.5-2026-04-23-xhigh, gpt-5.5-2026-04-23-medium, gpt-5.4-2026-03-05-medium, Claude Opus 4.7, Kimi K2.6, GLM-5.1, Qwen3.5-27B, Cursor, MiniMax M2.7, Gemini 3.1 Pro Preview, Claude Sonnet 4.6, Qwen3.5-397B-A17B, gpt-5.3-codex-xhigh, gpt-5.3-codex, Qwen3.5-35B-A3B, Claude Opus 4.6, GLM-5, MiniMax M2.5, Codex, Qwen3-Coder-Next, GLM-4.7 Flash, gpt-5.2-codex, gpt-5.2-2025-12-11-xhigh, gpt-5.1-codex, GLM-4.7, gpt-5-mini-2025-08-07-high, gpt-oss-120b-high, Kimi K2 Thinking, MiniMax M2.1, gpt-5.1-codex-max, gpt-5.2-2025-12-11-medium, Devstral-2-123B-Instruct-2512, Devstral-Small-2-24B-Instruct-2512, and DeepSeek-V3.2. These models represent the latest developments in AI for coding tasks.
Several models were deprecated from the leaderboard, including gpt-5.2-2025-12-11-xhigh, gpt-5.1-codex-max, gpt-5.1-codex, gpt-5-mini-2025-08-07-high, gpt-5-mini-2025-08-07-medium, Qwen3-235B-A22B-Instruct-2507, DeepSeek-R1-0528, Qwen3-Coder-30B-A3B-Instruct, Qwen3-Next-80B-A3B-Instruct, Qwen3-30B-A3B-Instruct-2507, gpt-5-2025-08-07-medium, gpt-5-2025-08-07-high, Claude Sonnet 4, Claude Opus 4.1, o3-2025-04-16, gpt-5-codex, GLM-4.5, o4-mini-2025-04-16, gpt-5-2025-08-07-minimal, gpt-4.1-2025-04-14, Qwen3-235B-A22B-Thinking-2507, gpt-4.1-mini-2025-04-14, Qwen3-30B-A3B-Thinking-2507, gemini-2.5-pro, gemini-2.5-flash, and DeepSeek-V3.1. This removal of older models indicates a focus on maintaining a current and relevant evaluation set.
The leaderboard evaluates AI models on 111 problems sourced from 65 repositories. Reference evaluations for Junie CLI and Claude Code were also added, providing benchmarks for comparison. These evaluations help assess the performance of various AI models in practical software engineering scenarios.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
A leaderboard tracking AI models for Software Engineering (SWE) tasks has seen several updates, with new models like GLM 5.2, DeepSeek-V4 Pro, and Gemini 3.5 Flash added, while older models such as gpt-5.2-2025-12-11-xhigh and Claude Sonnet 4 were deprecated. These changes reflect ongoing advancements and refinements in AI capabilities for code generation and software development assistance.