The UK's AISI (Artificial Intelligence Safety Institute) and EvalEval have announced a new phase in their collaboration, focusing on making AI benchmark results reproducible. This initiative builds on previous research and feedback that shaped the Every Eval Ever (EEE) schema, now being put into practice through shared infrastructure.
As AI deployment increases, model evaluations are critical for understanding performance. However, results are often reported in inconsistent formats across various platforms, lacking the necessary information for reproduction. This new collaboration aims to close these gaps by standardizing reporting and providing a common structure for evaluation results.
EvalEval's mission is to enhance the evaluation ecosystem using its Every Eval Ever reporting schema and the open platform, Evaluation Cards. AISI is now making publicly reported evaluation methods and findings available through Evaluation Cards. This release includes verified results, context, and configuration information for five benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0.
The data covers six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. Additionally, results from two cyber evaluations, Cyber CTFs and The Last Ones, are included. This data accompanies AISI's paper, "How Inference Compute Shapes Frontier LLM Evaluation," which examines how benchmark performance is influenced by inference-time compute and evaluation protocols.
This initiative is significant for the AI community as it promotes greater transparency and reproducibility in model evaluations. By standardizing how results are reported and providing detailed context, researchers and developers can better understand and compare AI model performance, fostering more rigorous and reliable AI development.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
The UK's AISI and EvalEval have released publicly reported evaluation methods and findings through EvalEval's Evaluation Cards platform, including verified results for five benchmarks and six frontier models. This collaboration aims to improve the reproducibility and transparency of AI model evaluations by providing a shared infrastructure for reporting and interpreting results.