ENT-ff3889bc-c073-43cd-b5a3-2b5aa262e589 · benchmark · evaluation framework revision
Holistic Evaluation of Language Models (HELM)
Reviewed identifier stanford-crfm/helm@63754d05db6f874e41a395880fb573890a13e791 · current status intentionally unknown
Scope
Reviewed benchmark or evaluation-protocol identity; model results and leaderboard ranks are separate, unrepresented observations.
Boundary
HELM entered maintenance mode on 2026-06-01. Framework availability is not evidence that every hosted leaderboard result is current or comparable.
Benchmark comparability passport
- Publisher
- Stanford CRFM
- Version
- source revision 63754d05; no semantic benchmark version asserted
- Task scope
- Framework for standardized, reproducible evaluation of language and multimodal foundation models across datasets and benchmarks.
- Metric contract
- Scenario-specific metrics include dimensions beyond accuracy, such as efficiency, bias and toxicity; no single HELM score is represented.
- Evaluation conditions
- A result is conditioned on the pinned framework revision, run entry, suite, model adapter, scenario and metric configuration.
- Comparability boundary
- Compare only runs with aligned HELM revision, scenario, adapter, prompting/run specification and metric; the `latest` route is not a fixed version.
- Contamination boundary
- The selected source does not establish absence of benchmark exposure in model training data.
This passport describes a benchmark definition or protocol. It publishes no model score, rank or cross-version equivalence.
Source-qualified relations
No reviewed relation is asserted for this identity.
Atlas claims with an exact source route
No exact source-URL connection in r0056. No name match was substituted.
Every displayed attribute links to source-qualified statement IDs. Relations are explicit and source-qualified. Identity does not imply ownership, legal status, present availability or equivalence with a mutable alias.