FrontierHarness Eval revealed a significant disparity in the cost per pass when evaluating AI models. Across nine different harnesses using the same underlying model, the cost per pass varied by a factor of 17. This indicates that the choice of evaluation harness can substantially influence the financial outlay for AI model testing.
The analysis noted that excluding failed attempts can misrepresent the true cost; including them increased the cost per task to $3.24. Additionally, cache hit rates do not directly equate to cost savings, as a cached failure can still incur higher costs than a short cache miss. One instance showed a harness passing 19 tasks but reaching $18.34 in cost per task, illustrating that quality and cost can diverge.
The evaluations were performed on Runta agent runtimes, ensuring consistent testing conditions. Each run started from a fresh restore with identical vCPU, memory, disk size, disk contents, and memory state. FrontierHarness v1.0 is specifically designed for software engineering contexts and terminal-based tasks, and its findings may not be generalizable to other areas of knowledge work.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
FrontierHarness Eval demonstrated that the cost per pass for evaluating AI models can vary by up to 17 times across different harnesses, even when using the same model. This variation highlights the impact of evaluation methodology on the efficiency and cost-effectiveness of AI development.