Today: the five AI headlines of the day → and the AI Wiki
What an entity record is How to read this page

A named actor, followed across desks

An organisation, model, chip or programme, with the statements attached to it, so that the same actor can be tracked without re-reading every domain.

Statements carry their own sources

Each statement on this page is bound to a source and a date. The entity page collects them; it does not average them into a verdict.

Absence is not evidence

An entity with few statements has been observed less, not necessarily done less. Read the count as coverage, never as significance.

New to this publication?

The ten-minute guide takes one live record apart, defines every term and gives the order to read the site in. Start here →

ENT-3a6be454-fedf-4237-8d71-f378e916b0a7 · benchmark · live human preference evaluation policy

Arena Leaderboard Policy

Reviewed identifier arena-leaderboard-policy-2026-09-01 · current status intentionally unknown

Scope

Reviewed benchmark or evaluation-protocol identity; model results and leaderboard ranks are separate, unrepresented observations.

Boundary

Human preference is not factual correctness, safety or universal capability. Model listing/deprecation and changing votes prevent a static rank from becoming an entity attribute.

Benchmark comparability passport

Publisher
Arena Team
Version
last updated 2026-09-01
Task scope
Live community evaluation of eligible foundation models through randomized battles and human preference votes across Arena surfaces.
Metric contract
Leaderboard ratings are estimated from battle outcomes with a regression that reweights sampling; ranks and intervals depend on the current methodology and accumulated vote data.
Evaluation conditions
Result identity includes Arena surface/category, policy and methodology revision, model endpoint identity and availability, battle/vote window, filters, sampling/reweighting and data-cleaning rules.
Comparability boundary
The leaderboard is live and methodology changes are logged; a rank is not timeless and must retain its observation time, category, policy/method version, interval and vote window.
Contamination boundary
Live user prompts reduce dependence on one fixed public test set, but the policy does not establish absence of manipulation, selection effects, provider exposure or preference bias.

This passport describes a benchmark definition or protocol. It publishes no model score, rank or cross-version equivalence.

Source-qualified relations

  1. No reviewed relation is asserted for this identity.

Atlas claims with an exact source route

  1. No exact source-URL connection in r0056. No name match was substituted.

Every displayed attribute links to source-qualified statement IDs. Relations are explicit and source-qualified. Identity does not imply ownership, legal status, present availability or equivalence with a mutable alias.