Today: the five AI headlines of the day → and the AI Wiki
What a signpost is How to read this page

A watched question with a line drawn in advance

The threshold was written before the observation existed. That is what stops a state from being argued into place after the fact.

Four states, no probabilities

Observed, emerging, contested, not observed. These are qualitative readings against the threshold. No number is invented here, because an invented number would be the most quotable and least true thing on the page.

It stays alive between revisions

The update condition names what would move the state. A signpost that has stayed quiet is information too, and is published rather than hidden.

New to this publication?

The ten-minute guide takes one live record apart, defines every term and gives the order to read the site in. Start here →

SCN-QUBE-003 · Observed 2025-02-18

contested

Deterministic Automation Backlash

Qualitative signpost review · no probability assigned

Bounded observation

OpenAI's SWE-Lancer evaluation used more than 1,400 real freelance software-engineering tasks with end-to-end tests and reported that frontier models were still unable to solve the majority. The result supports retaining deterministic baselines and bounded autonomy, but it does not show organizations actually rolling back autonomy tiers.

Trigger threshold

Move to observed for a workflow when representative production comparisons show deterministic or copilot-only systems outperform agent loops on accepted output after review time, defects, reversals and liability cost, followed by a documented autonomy rollback.

Counter-indicator

A benchmark failure rate is not a production cost comparison; updated models, scaffolds and workflow selection can change results, and the source records capability limits rather than an enterprise backlash.

Update criterion

Review updated SWE-Lancer results and controlled production evaluations by workflow; upgrade only with net-outcome superiority plus an operational rollback, and retire per workflow after repeatable bounded-agent superiority.

Primary routes

  1. https://openai.com/index/swe-lancer/