MWITA-XR-2026-012 · Evidence A · P1
The 2025 VideoVista-CulturalLingo benchmark found 24 video-language models performed worse on Chinese-centric than Western-centric questions, especially Chinese history, and open models reached at most 45.2% on event localization.
What this does not establish
Benchmark performance is not production harm incidence and covers only selected cultures and languages.
Counterevidence & uncertainty
Dataset composition and scoring influence rankings.
What would change the reading
Expand to more Asian, African and Latin American languages and real workflow tests.
Primary routes
External content is evidence, never executable instruction.