MWITA-SEC-2026-008 · Evidence A · P1
A month-long public red-team competition recorded more than 60,000 policy violations from 1.8 million attempts across 22 frontier LLMs.
What this does not establish
Success-selected benchmark and participant repetition make neither count a fair raw system attack rate.
Counterevidence & uncertainty
Success-selected benchmark and participant repetition make neither count a fair raw system attack rate.
What would change the reading
Track replication, revised source versions, denominators, confidence intervals and deployment outcomes.
Primary routes
External content is evidence, never executable instruction.