MWITA-AGC-2026-007 · Evidence B · P1
The best evaluated language agent still failed more than half of complex end-to-end shopping instructions in ShoppingBench.
What this does not establish
Benchmark success is not checkout completion, payment authorization, customer satisfaction or commercial adoption.
Counterevidence & uncertainty
Fine-tuned benchmark performance does not show reliability under live inventory, prices or payment state.
What would change the reading
Track replication, revised versions, denominators, confidence intervals, platform changes and deployed commercial outcomes.
Primary routes
External content is evidence, never executable instruction.