MWITA-CD-2026-066 · Evidence B · P0
Deepfake-Eval-2024 reports that evaluated open-source state-of-the-art detectors lost 50% video AUC, 48% audio AUC and 45% image AUC relative to older benchmarks on a 52-language in-the-wild dataset collected from 88 websites; commercial systems remained below forensic analysts.
What this does not establish
The reported drops do not quantify every commercial detector, future model or deployment threshold, and the preprint is not itself a regulatory performance standard.
Counterevidence & uncertainty
Collection, labeling, class balance and version choices affect generalization; the paper was revised through May 2026.
What would change the reading
Update on peer review, dataset audits, reproducible evaluation of deployed systems and newer in-the-wild distributions.
Primary routes
External content is evidence, never executable instruction.