MWITA-ME-2026-025 · Evidence A · P0
DialectalArabicMMLU manually adapted 3,000 question-answer pairs into each of five dialects—Syrian, Egyptian, Emirati, Saudi and Moroccan—across 32 domains and found substantial performance variation among 19 tested open-weight models.
What this does not establish
Modern Standard Arabic performance cannot be substituted for dialect performance, and the benchmark does not prove conversational or safety quality.
Counterevidence & uncertainty
Multiple choice, translation/adaptation choices, model sizes, prompt design and dialect continua constrain external validity.
What would change the reading
Evaluate the actual production model on native-authored target tasks, code-switching, safety, retrieval, latency and cost per country.
Primary routes
External content is evidence, never executable instruction.