MWITA-CA-2026-036 · Evidence B · P1
VoiceWukong was designed to compare humans, dedicated detectors and a multimodal language model in voice-deepfake detection scenarios.
What this does not establish
Benchmark ranking does not guarantee production performance or evidentiary reliability.
Counterevidence & uncertainty
Distribution shift and adversarial adaptation remain.
What would change the reading
Independent in-the-wild evaluation.
Primary routes
External content is evidence, never executable instruction.