MWITA-MIO-2026-002 · Evidence A · P1
In 663 Facebook RCT benchmarks, neither double/debiased machine learning nor stratified propensity-score matching reliably recovered causal ad lift despite more than 5,000 user- and experiment-level features.
What this does not establish
Does not prove nonexperimental measurement can never work; authors identify missing auction-specific features as a possible route, and platform/campaign selection bounds transport.
Counterevidence & uncertainty
Flexible deep-learning nuisance models did not cure the discrepancy, countering the assumption that modern ML plus rich user data substitutes for randomized incrementality tests.
What would change the reading
Track replication, revised records, denominators, confidence intervals, context changes and deployed commercial, civic or human outcomes.
Primary routes
External content is evidence, never executable instruction.