MWITA-CA-2026-032 · Evidence B · P1
RobustSora introduced a 6,500-video benchmark with authentic-clean, authentic-with-fake-watermark, generated-watermarked and generated-de-watermarked classes.
What this does not establish
A benchmark dataset is not representative prevalence or deployed field performance.
Counterevidence & uncertainty
Preprint and dataset construction choices limit generalization.
What would change the reading
Peer review and external replications.
Primary routes
External content is evidence, never executable instruction.