SCN-QUBE-003 · Observed 2025-02-18
contestedDeterministic Automation Backlash
Qualitative signpost review · no probability assigned
Bounded observation
OpenAI's SWE-Lancer evaluation used more than 1,400 real freelance software-engineering tasks with end-to-end tests and reported that frontier models were still unable to solve the majority. The result supports retaining deterministic baselines and bounded autonomy, but it does not show organizations actually rolling back autonomy tiers.
Trigger threshold
Move to observed for a workflow when representative production comparisons show deterministic or copilot-only systems outperform agent loops on accepted output after review time, defects, reversals and liability cost, followed by a documented autonomy rollback.
Counter-indicator
A benchmark failure rate is not a production cost comparison; updated models, scaffolds and workflow selection can change results, and the source records capability limits rather than an enterprise backlash.
Update criterion
Review updated SWE-Lancer results and controlled production evaluations by workflow; upgrade only with net-outcome superiority plus an operational rollback, and retire per workflow after repeatable bounded-agent superiority.