0 numerical observations admitted · 9 reviewed benchmark passports
Schema before scores.
A benchmark result becomes publishable only as an immutable, source-receipted observation with explicit system identity, protocol, metric, uncertainty, contamination, verification and correction history.
DIRECT
Both observations pass the admission contract and every comparison-key dimension is exactly aligned.
CONDITIONED
A disclosed, source-supported transformation permits a qualified comparison; assumptions and residual differences remain visible.
CONTEXT_ONLY
The records may be juxtaposed as separate facts, but no quantitative superiority or trend inference is permitted.
BLOCKED
Identity, rights, integrity, verification or validity defects prohibit comparison publication.
No current score or rank is asserted.
No current score, rank or model-performance claim is admitted until an immutable observation passes schema validation, exact entity resolution, source-receipt verification, rights review and the comparison gate.
Reviewed benchmark definitions
These passports define measurement objects and their boundaries. They are not model-performance observations.
- Reviewed benchmark passportArena Leaderboard Policylast updated 2026-09-01
- Reviewed benchmark passportBIG-benchsource revision 092b196c; no semantic benchmark version stated
- Reviewed benchmark passportGPQAsource revision 56686c06; no semantic benchmark version stated
- Reviewed benchmark passportHolistic Evaluation of Language Models (HELM)source revision 63754d05; no semantic benchmark version asserted
- Reviewed benchmark passportHumanEvalsource revision 6d43fb98; no semantic benchmark version stated
- Reviewed benchmark passportMassive Multitask Language Understanding (MMLU)source revision 4450500f; no semantic benchmark version stated
- Reviewed benchmark passportMETR Task-Completion Time Horizon 1.11.1
- Reviewed benchmark passportMLPerf Inference v6.06.0
- Reviewed benchmark passportSWE-bench VerifiedVerified subset; source revision 02e7a74f
Publication is not comparability. Equal metadata does not prove causal validity, representativeness or deployment fitness.