← All stories
● Covered by 1 source · 1 reportLow impact1 neutral

Big Pickle Model Achieves 50.8% Task Resolve Rate on SWE Atlas Codebase QnA Benchmark

🔄 Updated 7h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Big Pickle model scored 50.8% on SWE Atlas Codebase QnA.
  • It outperformed all mini-swe-agent scaffold entries.
  • The model is currently free during its stealth period.
  • Evaluation followed Scale's published protocol with some resource adjustments.

Big Pickle's Benchmark Performance

The Big Pickle model, developed by OpenCode Zen, recorded a 50.8% task resolve rate on the SWE Atlas Codebase QnA benchmark. This evaluation involved 124 tasks and utilized the mini-swe-agent scaffold, a minimal bash-only scaffold used for non-first-party models on the leaderboard.

Comparison with Existing Models

Within its scaffold class, Big Pickle surpassed all other entries on the official SWE Atlas QnA leaderboard. It also outperformed GPT models running on the Codex-scaffold. Only two Claude models, operating on their native Claude Code scaffold, achieved higher scores.

Evaluation Methodology and Caveats

The evaluation adhered to Scale AI's published protocol, including the use of Harbor v0.18.0 with Modal sandboxes and the specified judge model, claude-opus-4-5-20251101. However, the run used a single trial per task instead of the official three, and reduced sandbox resources (4 CPU / 8 GB) compared to the declared 16 CPU / 16 GB. These adjustments were made to fit a personal budget and are noted as potential factors that could depress, but not inflate, the score. The evaluation was self-reported and not verified by Scale AI.

Model Availability and Cost

The Big Pickle model is currently available for free during its stealth period via OpenCode Zen's OpenAI-compatible endpoint. The total token consumption for the benchmark run was 674 million input tokens and 4.3 million output tokens, incurring no cost.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~7 min · 6 stories · Aug 15

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

The Big Pickle model, a free stealth model from OpenCode Zen, achieved a 50.8% task resolve rate on Scale AI's SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold. This performance outscored all other entries in the mini-swe-agent scaffold class and most GPT models, indicating its capability in code-related question answering tasks.