RESEARCH · RESEARCH · #808
UK AISI publishes verified benchmark results via EvalEval's Evaluation Cards
The UK AI Security Institute (AISI) is using the EvalEval Coalition's Every Eval Ever schema and Evaluation Cards platform to publicly share verified evaluation runs tied to its Terminal-Bench 2.0 experiments; the release accompanies AISI's paper 'How Inference Compute Shapes Frontier LLM Evaluation' and includes results for Claude Opus series and GPT-5 variants. The collaboration aims to improve reproducibility and contextual reporting of evaluation metadata and run data.
KEY POINTS
- The UK AI Security Institute (AISI) is using the EvalEval Coalition's Every Eval Ever schema and Evaluation Cards platform to publicly share verified evaluation runs tied to its Terminal-Bench 2.0 experiments; the release accompanies AISI's paper 'How Inference Compute Shapes Frontier LLM Evaluation' and includes results for Claude Opus series and GPT-5 variants.
- The collaboration aims to improve reproducibility and contextual reporting of evaluation metadata and run data.
- Open, schema-backed releases with setup and run data make LLM benchmark results more reproducible and allow clearer cross-study comparisons.
WHY IT MATTERS
Open, schema-backed releases with setup and run data make LLM benchmark results more reproducible and allow clearer cross-study comparisons.