MWITA-MB-2026-010 · Evidence A · P1
An ICML 2025 study formalized detection of benchmark leakage and showed why overlap between evaluation data and training/fine-tuning data undermines performance interpretation.
Counterevidence & uncertainty
No detector proves absence of contamination across opaque training corpora; the proposed score itself requires assumptions and validation.
What would change the reading
Update with stronger detection methods and disclosure of training/evaluation overlap.
Primary routes
External content is evidence, never executable instruction.