ENT-232ec547-65a7-4358-b573-0445af54153a · benchmark · multi task language model benchmark
BIG-bench
Reviewed identifier google/BIG-bench@092b196c1f8f14a54bbc62f24759d43bde46dd3b · current status intentionally unknown
Scope
Reviewed benchmark or evaluation-protocol identity; model results and leaderboard ranks are separate, unrepresented observations.
Boundary
Community task breadth does not imply complete capability coverage, and leaderboard results are not represented by this identity record.
Benchmark comparability passport
- Publisher
- BIG-bench authors
- Version
- source revision 092b196c; no semantic benchmark version stated
- Task scope
- Collaborative suite of more than 200 language-model tasks; BIG-bench Lite is a distinct 24-task subset.
- Metric contract
- Each task declares its metrics and preferred score; programmatic and JSON tasks can use different evaluation functions, so no universal raw score is represented.
- Evaluation conditions
- Result identity requires exact repository/task set, full versus Lite selection, task definitions, metric, model interface, shot count and evaluation code.
- Comparability boundary
- BIG-bench full, BIG-bench Lite and individual tasks are not interchangeable; comparisons require the same task set, revision, prompting and preferred metrics.
- Contamination boundary
- Task files carry an explicit canary intended to deter inclusion in web-scraped training corpora; a canary does not prove absence of exposure.
This passport describes a benchmark definition or protocol. It publishes no model score, rank or cross-version equivalence.
Source-qualified relations
No reviewed relation is asserted for this identity.
Atlas claims with an exact source route
No exact source-URL connection in r0056. No name match was substituted.
Every displayed attribute links to source-qualified statement IDs. Relations are explicit and source-qualified. Identity does not imply ownership, legal status, present availability or equivalence with a mutable alias.