MWITA-AW-2026-016 · Evidence B · P1
A roughly four-month Singapore-government/Google sandbox tested computer-use agents in government-site QA, multilingual AI-safety testing, and social-assistance navigation; the published page reports successful seeded-defect detection and scalable safety testing but says implementation was not entirely error-free.
What this does not establish
It does not prove production reliability, labor savings, benefit eligibility accuracy, or superiority to scripted tools.
Counterevidence & uncertainty
No sample sizes, error rates or independent replication are provided on the page.
What would change the reading
Require task sets, baselines, human-review load, failures, reversibility and production follow-up.
Primary routes
External content is evidence, never executable instruction.