The Big Pickle model, developed by OpenCode Zen, recorded a 50.8% task resolve rate on the SWE Atlas Codebase QnA benchmark. This evaluation involved 124 tasks and utilized the mini-swe-agent scaffold, a minimal bash-only scaffold used for non-first-party models on the leaderboard.
Within its scaffold class, Big Pickle surpassed all other entries on the official SWE Atlas QnA leaderboard. It also outperformed GPT models running on the Codex-scaffold. Only two Claude models, operating on their native Claude Code scaffold, achieved higher scores.
The evaluation adhered to Scale AI's published protocol, including the use of Harbor v0.18.0 with Modal sandboxes and the specified judge model, claude-opus-4-5-20251101. However, the run used a single trial per task instead of the official three, and reduced sandbox resources (4 CPU / 8 GB) compared to the declared 16 CPU / 16 GB. These adjustments were made to fit a personal budget and are noted as potential factors that could depress, but not inflate, the score. The evaluation was self-reported and not verified by Scale AI.
The Big Pickle model is currently available for free during its stealth period via OpenCode Zen's OpenAI-compatible endpoint. The total token consumption for the benchmark run was 674 million input tokens and 4.3 million output tokens, incurring no cost.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
The Big Pickle model, a free stealth model from OpenCode Zen, achieved a 50.8% task resolve rate on Scale AI's SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold. This performance outscored all other entries in the mini-swe-agent scaffold class and most GPT models, indicating its capability in code-related question answering tasks.