How To Make AI Benchmark Results Reproducible With UK AISI And EvalEval
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How To Make AI Benchmark Results Reproducible With UK AISI And EvalEval on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute is publishing selected results from its evaluation work through EvalEval’s Evaluation Cards. The release covers five benchmarks across six frontier models, plus two cyber evaluations with a partly overlapping model set; the records add information about verification and evaluation setup.

The UK AI Security Institute (AISI) is publishing selected AI evaluation results through EvalEval’s Evaluation Cards, pairing scores with verification, evaluation context and configuration details, as described in the original analysis. The release accompanies AISI’s paper on how inference-time compute and evaluation protocols shape results, and includes five benchmarks tested across six frontier models, along with two cyber evaluations using a different, partly overlapping model set.

The paper’s main experiment covers HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The six models listed for those results are Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI also shared results from Cyber CTFs and The Last Ones, but the announcement says those evaluations use a different model set that overlaps only partly with the main experiment.

EvalEval says the Evaluation Cards contain verified results, evaluation context and configuration information. Its platform organizes benchmark metadata, evaluation-run data and model metadata in a shared format. The release is associated with AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how scores vary with inference-time compute and evaluation protocols.

For Humanity’s Last Exam, the paper tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased. The announcement says public methods and findings are being shared where appropriate; it does not describe the material as a complete archive of AISI’s evaluations.

At a glance
announcementWhen: Announced in connection with AISI’s pap…
The developmentAISI has begun sharing selected evaluation results through EvalEval’s Evaluation Cards, alongside details intended to help readers interpret how the results were produced.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Evaluation Setup Changes Scores

Benchmark scores can look directly comparable while reflecting different conditions. AISI’s Humanity’s Last Exam analysis illustrates that inference compute and feedback between attempts can affect how many tasks a model solves. Without information about those conditions, readers may not know what a reported score represents or whether two results can be fairly compared.

Putting selected results beside their setup details gives researchers and practitioners more information to inspect individual runs and compare them with other published evaluations. That can matter when evaluations inform model development, research and policy. The cards do not establish which benchmark or protocol is best, and the release alone does not show that every result can be reproduced independently. It makes some of the conditions behind selected results easier to examine.

A Shared Format for Evaluation Records

The collaboration follows earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from AISI helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that infrastructure to publicly reported AISI methods and findings.

AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas including transcript analysis and capability elicitation. EvalEval’s Evaluation Cards bring together results with benchmark and model information. The stated aim is to address a reporting problem: results released in different formats may leave out details needed to interpret a run, while repeating costly evaluations may not be feasible.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

What the Published Records Cover

The announcement does not specify the total number of records or transcripts, which setup fields are available for every benchmark, or whether independent researchers have reproduced the results. It also does not enumerate the models used in the two cyber evaluations, so readers should not assume the six-model list for the main experiment applies to them.

AISI says methods and findings are shared where appropriate, which leaves the release’s full coverage unclear. The announcement gives no record-by-record publication dates or process for resolving differences between results collected under different protocols. It also does not claim that every AISI evaluation or underlying transcript is included.

Broader Use of the EEE Schema

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. Under the EEE schema, model developers can submit verified results, while evaluation developers can report benchmarks and run data in a common format. Researchers in evaluation, governance and policy can explore the published cards by benchmark or model.

No further release date or adoption milestone was announced. Wider use could make comparisons across studies easier, but how useful the collection becomes will depend on the consistency and completeness of the records contributors publish.

Key Questions

Which benchmarks are included in AISI’s main experiment?

The paper’s main experiment reports results for HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.

Which models are listed for those five benchmarks?

The listed models are Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The cyber evaluations use a different, partly overlapping model set.

What information do the Evaluation Cards add?

EvalEval describes the cards as including verified results, evaluation context and configuration information, alongside benchmark and model metadata.

Do the published cards prove that results are independently reproducible?

No. The release is intended to support more reproducible and verifiable evaluation science, but the announcement does not say outside researchers have independently reproduced the results.

Is this a complete archive of AISI’s evaluations?

The announcement says publicly reported methods and findings are shared where appropriate. It does not claim that every AISI evaluation or underlying transcript is included.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

LinkedIn recruitment spam becomes Olde English prose after user hides AI prompt injection in bio — bots also also manipulated to address user as ‘My Lord’

A LinkedIn user injected an Old English prompt into their profile, causing recruiters’ messages to appear in 900 AD language, highlighting AI manipulation risks.

Discover The 14 Best AI Automation Software For Streamlined Work In 2026

A 2026 ranking compares 14 AI automation guides, placing OpenCode Custom Workflows first while exposing limits in the available evidence.

The Newcomer Kimi K3 Takes #3 Spot On VigilSAR’s LLM Leaderboard

Moonshot’s Kimi K3 entered VigilSAR’s defense-ISR benchmark at No. 3, outperforming every listed GPT and Gemini model.

AI in Customer Support: Where It Helps and Where It Fails

For insights into AI’s role in customer support, discover where it excels and where human touch remains essential to truly meet customer needs.