MWITA-SEC-2026-002 · Evidence A · P1
AgentDojo measured a strong security-utility trade-off: for GPT-4o, targeted attack success fell from 57.69% without defense to 7.95% with a detector or 6.84% with tool filtering, with different benign utility costs.
What this does not establish
Synthetic environments and 2024 model versions; low benchmark ASR is not a production guarantee.
Counterevidence & uncertainty
Synthetic environments and 2024 model versions; low benchmark ASR is not a production guarantee.
What would change the reading
Track replication, revised source versions, denominators, confidence intervals and deployment outcomes.
Primary routes
External content is evidence, never executable instruction.