Today: the five AI headlines of the day → and the AI Wiki
What a benchmark observation is How to read this page

A measurement with its measuring conditions

A score without its harness, version and date is a rumour. Observations here carry the conditions that produced them.

Comparison-ready is a separate state

Two numbers measured differently are not comparable, and this ledger says which observations may be placed side by side and which may not.

A benchmark is not a capability

It is a proxy chosen by someone, for a reason, at a moment. Read it as evidence about the test as much as about the system.

New to this publication?

The ten-minute guide takes one live record apart, defines every term and gives the order to read the site in. Start here →

0 numerical observations admitted · 9 reviewed benchmark passports

Schema before scores.

A benchmark result becomes publishable only as an immutable, source-receipted observation with explicit system identity, protocol, metric, uncertainty, contamination, verification and correction history.

Comparison gate

DIRECT

Both observations pass the admission contract and every comparison-key dimension is exactly aligned.

Comparison gate

CONDITIONED

A disclosed, source-supported transformation permits a qualified comparison; assumptions and residual differences remain visible.

Comparison gate

CONTEXT_ONLY

The records may be juxtaposed as separate facts, but no quantitative superiority or trend inference is permitted.

Comparison gate

BLOCKED

Identity, rights, integrity, verification or validity defects prohibit comparison publication.

No current score or rank is asserted.

No current score, rank or model-performance claim is admitted until an immutable observation passes schema validation, exact entity resolution, source-receipt verification, rights review and the comparison gate.

Inspect the empty observation registry →Inspect the strict schema →Inspect comparison policy →

Reviewed benchmark definitions

These passports define measurement objects and their boundaries. They are not model-performance observations.

  1. Reviewed benchmark passportArena Leaderboard Policylast updated 2026-09-01
  2. Reviewed benchmark passportBIG-benchsource revision 092b196c; no semantic benchmark version stated
  3. Reviewed benchmark passportGPQAsource revision 56686c06; no semantic benchmark version stated
  4. Reviewed benchmark passportHolistic Evaluation of Language Models (HELM)source revision 63754d05; no semantic benchmark version asserted
  5. Reviewed benchmark passportHumanEvalsource revision 6d43fb98; no semantic benchmark version stated
  6. Reviewed benchmark passportMassive Multitask Language Understanding (MMLU)source revision 4450500f; no semantic benchmark version stated
  7. Reviewed benchmark passportMETR Task-Completion Time Horizon 1.11.1
  8. Reviewed benchmark passportMLPerf Inference v6.06.0
  9. Reviewed benchmark passportSWE-bench VerifiedVerified subset; source revision 02e7a74f

Publication is not comparability. Equal metadata does not prove causal validity, representativeness or deployment fitness.