MWITA-AF-2026-009 · Evidence A · P1
AfriHuBERT reports continued self-supervised pretraining on more than 10,000 hours spanning 1,226 African languages or dialects plus four widely used non-African languages, with gains on FLEURS language identification and speech recognition relative to mHuBERT-147.
What this does not establish
The source does not establish safe, accurate transcription for every speaker, dialect, acoustic setting, or domain.
Counterevidence & uncertainty
Aggregate metrics can hide per-language regressions; dataset licensing, speaker representation, noise robustness, and code-switching need separate review.
What would change the reading
Run per-language, per-accent, gender, noise, code-switch, safety, and latency tests on target-market audio with licensed data.
Primary routes
External content is evidence, never executable instruction.