MWITA-SSCO-2026-010 · Evidence B · P1
A Chinese sign-language-avatar benchmark showed measurable improvement but limited absolute quality: only four of ten submitted systems achieved good readability.
What this does not establish
A benchmark score is not everyday accessibility, task completion, or comprehension and covers only Chinese Sign Language.
Counterevidence & uncertainty
Relative benchmark progress coexists with modest absolute quality and incomplete evaluator metadata.
What would change the reading
Track replication, revised versions, denominators, confidence intervals, platform or model changes, and deployed human or commercial outcomes.
Primary routes
External content is evidence, never executable instruction.