MWITA-ME-2026-027 · Evidence A · P1
Atlas-Chat reports that its 9B Moroccan-Darija-adapted model outperformed a larger 13B comparator by 13% on the paper's DarijaMMLU suite and improved both discriminative and generative tasks.
What this does not establish
Smaller adapted models are not always better, and benchmark improvement does not establish business or user outcomes.
Counterevidence & uncertainty
Training-data representativeness, licensing, contamination, evaluation construction and code-switching affect generalization.
What would change the reading
Compare current models on native-authored Moroccan tasks, Amazigh/Arabic/French code-switching, safety, latency and human preference.
Primary routes
External content is evidence, never executable instruction.