ENT-da6a449a-c99c-42d4-8e13-f5b1280fd72e · benchmark · agent reliability metric suite
METR Task-Completion Time Horizon 1.1
Reviewed identifier METR Time Horizon 1.1 · current status intentionally unknown
Scope
Reviewed benchmark or evaluation-protocol identity; model results and leaderboard ranks are separate, unrepresented observations.
Boundary
Time horizon is task difficulty measured in human time, not agent wall-clock autonomy, coverage of all intellectual work, job automation or deployment reliability.
Benchmark comparability passport
- Publisher
- METR
- Version
- 1.1
- Task scope
- More than one hundred self-contained software-engineering, machine-learning and cybersecurity tasks with human expert duration estimates.
- Metric contract
- Predicted human-expert task duration where an agent reaches a specified reliability, derived from a logistic success curve; the page reports 50% and 80% horizons.
- Evaluation conditions
- Result identity includes model, scaffold, task-suite version, human-duration method, test split, reliability threshold, token/time limits, six independent task runs and any human re-scoring.
- Comparability boundary
- Compare only aligned suite version, task distribution, scaffold, duration model and reliability threshold; values above sixteen hours are marked unreliable for the current suite.
- Contamination boundary
- Some tasks are private, but the source does not prove that every evaluated model lacked exposure to every task or analogous task.
This passport describes a benchmark definition or protocol. It publishes no model score, rank or cross-version equivalence.
Source-qualified relations
No reviewed relation is asserted for this identity.
Atlas claims with an exact source route
Every displayed attribute links to source-qualified statement IDs. Relations are explicit and source-qualified. Identity does not imply ownership, legal status, present availability or equivalence with a mutable alias.