MWITA-MB-2026-008 · Evidence A · P1
τ-bench introduced pass^k to measure the probability that an agent succeeds consistently across repeated tool-user interactions, exposing reliability that single-attempt success can hide.
Counterevidence & uncertainty
Users are simulated, policies and APIs are simplified, and benchmark tasks may not match production distributions.
What would change the reading
Update with production-calibrated tasks, independent reruns and contamination checks.
Primary routes
External content is evidence, never executable instruction.