Today: the five AI headlines of the day → and the AI Wiki
What an entity record is How to read this page

A named actor, followed across desks

An organisation, model, chip or programme, with the statements attached to it, so that the same actor can be tracked without re-reading every domain.

Statements carry their own sources

Each statement on this page is bound to a source and a date. The entity page collects them; it does not average them into a verdict.

Absence is not evidence

An entity with few statements has been observed less, not necessarily done less. Read the count as coverage, never as significance.

New to this publication?

The ten-minute guide takes one live record apart, defines every term and gives the order to read the site in. Start here →

ENT-da6a449a-c99c-42d4-8e13-f5b1280fd72e · benchmark · agent reliability metric suite

METR Task-Completion Time Horizon 1.1

Reviewed identifier METR Time Horizon 1.1 · current status intentionally unknown

Scope

Reviewed benchmark or evaluation-protocol identity; model results and leaderboard ranks are separate, unrepresented observations.

Boundary

Time horizon is task difficulty measured in human time, not agent wall-clock autonomy, coverage of all intellectual work, job automation or deployment reliability.

Benchmark comparability passport

Publisher
METR
Version
1.1
Task scope
More than one hundred self-contained software-engineering, machine-learning and cybersecurity tasks with human expert duration estimates.
Metric contract
Predicted human-expert task duration where an agent reaches a specified reliability, derived from a logistic success curve; the page reports 50% and 80% horizons.
Evaluation conditions
Result identity includes model, scaffold, task-suite version, human-duration method, test split, reliability threshold, token/time limits, six independent task runs and any human re-scoring.
Comparability boundary
Compare only aligned suite version, task distribution, scaffold, duration model and reliability threshold; values above sixteen hours are marked unreliable for the current suite.
Contamination boundary
Some tasks are private, but the source does not prove that every evaluated model lacked exposure to every task or analogous task.

This passport describes a benchmark definition or protocol. It publishes no model score, rank or cross-version equivalence.

Source-qualified relations

  1. No reviewed relation is asserted for this identity.

Atlas claims with an exact source route

  1. MWITA-MB-2026-007METR defines task-completion time horizon as the human-duration point at which an agent is predicted to succeed at a specified reliability under a particular task suite and scaffold.Exact source-URL connection

Every displayed attribute links to source-qualified statement IDs. Relations are explicit and source-qualified. Identity does not imply ownership, legal status, present availability or equivalence with a mutable alias.