MWITA-ME-2026-026 · Evidence A · P1
TounsiBench evaluated ten widely used LLMs on 744 Tunisian-Arabic instructions with human-written references and found that most tested models struggled to recognize and respond in Tunisian Arabic.
What this does not establish
The result does not describe all Tunisian speakers, domains or later models and does not validate LLM-as-judge alone.
Counterevidence & uncertainty
Automated final leaderboard depends on GPT-4o as judge despite reported human correlation; dataset size and instruction mix constrain coverage.
What would change the reading
Retest production systems with native reviewers across region, age, domain, code-switching, safety and real conversations.
Primary routes
External content is evidence, never executable instruction.