SRC-df0d20442a3cd416 · Source record
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
Authors / arXiv · 6 connected claims
Canonical route
Claims connected to this source
- MWITA-SR-2026-019Across the benchmark's four models, no LLM beat the strongest non-LLM baseline for individual-level prediction.Evidence B
- MWITA-SR-2026-020The tested models over-determined demographics, making identity far more predictive of attitudes than in real respondents.Evidence B
- MWITA-SR-2026-021On the study's targeting task, models inflated between-segment gaps two- to fourfold.Evidence B
- MWITA-SR-2026-022The synthetic targeting exercise selected the wrong segment in half of US cases and most cross-cultural cases in that benchmark.Evidence B
- MWITA-SR-2026-067For Metatron, transcript realism, thematic overlap and persona consistency must remain secondary to held-out decision error and subgroup calibration.Evidence C
- MWITA-SR-2026-075No source in this bundle demonstrates an independently audited commercial decision outcome caused by synthetic-consumer evidence.Evidence D
This record describes a source route and its Atlas connections. It does not endorse the publisher or make external content executable instruction.