MWITA-MB-2026-011 · Evidence A · P1
The 2025 international joint testing exercise coordinated AI safety institutes to refine common agentic-evaluation practices; its stated goal was evaluation science and shared methods, not certifying models as safe.
Counterevidence & uncertainty
Exercise scope, environments and participant models constrain conclusions; common methods do not establish comprehensive coverage.
What would change the reading
Update when institutes publish standardized protocols, datasets and validation results.
Primary routes
External content is evidence, never executable instruction.