Today: the five AI headlines of the day → and the AI Wiki
What an entity record is How to read this page

A named actor, followed across desks

An organisation, model, chip or programme, with the statements attached to it, so that the same actor can be tracked without re-reading every domain.

Statements carry their own sources

Each statement on this page is bound to a source and a date. The entity page collects them; it does not average them into a verdict.

Absence is not evidence

An entity with few statements has been observed less, not necessarily done less. Read the count as coverage, never as significance.

New to this publication?

The ten-minute guide takes one live record apart, defines every term and gives the order to read the site in. Start here →

ENT-3e5d2087-d0b9-4f35-b4a7-dcd41c022197 · benchmark · code generation functional correctness benchmark

HumanEval

Reviewed identifier openai/human-eval@6d43fb980f9fee3c892a914eda09951f772ad10d · current status intentionally unknown

Scope

Reviewed benchmark or evaluation-protocol identity; model results and leaderboard ranks are separate, unrepresented observations.

Boundary

Executing generated code is explicitly unsafe without a robust sandbox. Functional-test passage is not proof of secure, maintainable or production-ready code.

Benchmark comparability passport

Publisher
OpenAI
Version
source revision 6d43fb98; no semantic benchmark version stated
Task scope
Hand-written problem-solving dataset for generated Python function completions evaluated by executable tests.
Metric contract
pass@k estimated from multiple generated samples; the official evaluator refuses cases with fewer samples than k because no unbiased estimator is available.
Evaluation conditions
Result identity includes problem file, sample count per task, k values, generation settings, completion format, evaluator revision and sandbox/runtime resources.
Comparability boundary
Compare only results with aligned task data, sampling count and generation setup, k, evaluator revision and runtime; pass@1 and pass@100 answer different questions.
Contamination boundary
The selected source does not establish absence of HumanEval problem or solution exposure in training data.

This passport describes a benchmark definition or protocol. It publishes no model score, rank or cross-version equivalence.

Source-qualified relations

  1. No reviewed relation is asserted for this identity.

Atlas claims with an exact source route

  1. No exact source-URL connection in r0056. No name match was substituted.

Every displayed attribute links to source-qualified statement IDs. Relations are explicit and source-qualified. Identity does not imply ownership, legal status, present availability or equivalence with a mutable alias.