🔍 Read the full analysis: How To Make AI Benchmark Results Reproducible With UK AISI And EvalEval on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
The UK AI Security Institute is publishing selected results from its evaluation work through EvalEval’s Evaluation Cards. The release covers five benchmarks across six frontier models, plus two cyber evaluations with a partly overlapping model set; the records add information about verification and evaluation setup.
The paper’s main experiment covers HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The six models listed for those results are Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI also shared results from Cyber CTFs and The Last Ones, but the announcement says those evaluations use a different model set that overlaps only partly with the main experiment.
EvalEval says the Evaluation Cards contain verified results, evaluation context and configuration information. Its platform organizes benchmark metadata, evaluation-run data and model metadata in a shared format. The release is associated with AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how scores vary with inference-time compute and evaluation protocols.
For Humanity’s Last Exam, the paper tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased. The announcement says public methods and findings are being shared where appropriate; it does not describe the material as a complete archive of AISI’s evaluations.
Why Evaluation Setup Changes Scores
Benchmark scores can look directly comparable while reflecting different conditions. AISI’s Humanity’s Last Exam analysis illustrates that inference compute and feedback between attempts can affect how many tasks a model solves. Without information about those conditions, readers may not know what a reported score represents or whether two results can be fairly compared.
Putting selected results beside their setup details gives researchers and practitioners more information to inspect individual runs and compare them with other published evaluations. That can matter when evaluations inform model development, research and policy. The cards do not establish which benchmark or protocol is best, and the release alone does not show that every result can be reproduced independently. It makes some of the conditions behind selected results easier to examine.
The collaboration follows earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from AISI helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that infrastructure to publicly reported AISI methods and findings.
AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas including transcript analysis and capability elicitation. EvalEval’s Evaluation Cards bring together results with benchmark and model information. The stated aim is to address a reporting problem: results released in different formats may leave out details needed to interpret a run, while repeating costly evaluations may not be feasible.
“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”
— EvalEval Coalition
What the Published Records Cover
The announcement does not specify the total number of records or transcripts, which setup fields are available for every benchmark, or whether independent researchers have reproduced the results. It also does not enumerate the models used in the two cyber evaluations, so readers should not assume the six-model list for the main experiment applies to them.
AISI says methods and findings are shared where appropriate, which leaves the release’s full coverage unclear. The announcement gives no record-by-record publication dates or process for resolving differences between results collected under different protocols. It also does not claim that every AISI evaluation or underlying transcript is included.
Broader Use of the EEE Schema
EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. Under the EEE schema, model developers can submit verified results, while evaluation developers can report benchmarks and run data in a common format. Researchers in evaluation, governance and policy can explore the published cards by benchmark or model.
No further release date or adoption milestone was announced. Wider use could make comparisons across studies easier, but how useful the collection becomes will depend on the consistency and completeness of the records contributors publish.
Key Questions
Which benchmarks are included in AISI’s main experiment?
The paper’s main experiment reports results for HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.
Which models are listed for those five benchmarks?
The listed models are Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The cyber evaluations use a different, partly overlapping model set.
What information do the Evaluation Cards add?
EvalEval describes the cards as including verified results, evaluation context and configuration information, alongside benchmark and model metadata.
Do the published cards prove that results are independently reproducible?
No. The release is intended to support more reproducible and verifiable evaluation science, but the announcement does not say outside researchers have independently reproduced the results.
Is this a complete archive of AISI’s evaluations?
The announcement says publicly reported methods and findings are shared where appropriate. It does not claim that every AISI evaluation or underlying transcript is included.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
