MWITA-AF-2026-008 · Evidence A · P0
IrokoBench's peer-reviewed evaluation of 17 human-translated low-resource African languages found the best tested open model, Gemma 2 27B, reached 63% of the best tested proprietary model GPT-4o across its benchmark, while translate-to-English improved some larger English-centric models.
What this does not establish
The result does not show that proprietary models are universally better or that English translation preserves meaning in operational settings.
Counterevidence & uncertainty
Benchmark contamination, translation artifacts, prompt sensitivity, closed-model drift, and task selection constrain external validity.
What would change the reading
Re-evaluate current candidate models on locally reviewed task sets, dialects, code-switching, safety, latency, and cost before launch.
Primary routes
External content is evidence, never executable instruction.