ENT-9a3c06c8-568b-4f39-9f62-e1edb55e8a6a · benchmark · graduate science multiple choice benchmark
GPQA
Reviewed identifier idavidrein/gpqa@56686c06f5e19865c153de0fdb11be3890014df7 · current status intentionally unknown
Scope
Reviewed benchmark or evaluation-protocol identity; model results and leaderboard ranks are separate, unrepresented observations.
Boundary
The `Google-Proof` title is the benchmark name, not a guarantee that every item is unsearchable, uncontaminated or representative of all expert reasoning.
Benchmark comparability passport
- Publisher
- GPQA authors
- Version
- source revision 56686c06; no semantic benchmark version stated
- Task scope
- Graduate-level, Google-proof question-answering dataset with separate data files and baseline implementations.
- Metric contract
- A result is meaningful only with the exact data file/subset, shuffled answer-choice seed and prompt/retrieval mode; no single aggregate is admitted here.
- Evaluation conditions
- Result identity includes repository/data revision, subset, model, zero/few-shot or chain-of-thought mode, closed/open-book retrieval setting, answer-choice shuffle seed and cache behavior.
- Comparability boundary
- Do not combine GPQA subsets or compare retrieval and closed-book runs, different prompts, seeds or code revisions as one measurement.
- Contamination boundary
- The dataset includes a canary string; this is a detection/deterrence aid, not proof that training exposure did not occur.
This passport describes a benchmark definition or protocol. It publishes no model score, rank or cross-version equivalence.
Source-qualified relations
No reviewed relation is asserted for this identity.
Atlas claims with an exact source route
No exact source-URL connection in r0056. No name match was substituted.
Every displayed attribute links to source-qualified statement IDs. Relations are explicit and source-qualified. Identity does not imply ownership, legal status, present availability or equivalence with a mutable alias.