MWITA-MR-2026-011 · Evidence A · P1
A 2025 human-versus-GPT-4 study across Saudi Arabia, UAE and the US found reasonable aggregate alignment but generally weaker non-WEIRD correlations, systematic country/domain biases, and poor prediction of which human treatment effects were significant.
Counterevidence & uncertainty
Samples were online and nonrepresentative and covered three policy domains, not product purchase; newer models may differ.
What would change the reading
Representative preregistered replications across products/languages achieve stable item-, segment- and treatment-effect accuracy.
Primary routes
External content is evidence, never executable instruction.