MWITA-EL-2026-016 · Evidence B · P1
A 2025 study of 300 academic-English placement essays found GPT-4 scoring had high within-model reliability and moderate positive correlation with human scores; detailed rubrics, rationales, examples, linguistic features and averaging multiple ratings improved alignment.
What this does not establish
Reliability and local placement agreement do not establish construct validity, fairness or suitability for other prompts, languages, models or high-stakes decisions.
Counterevidence & uncertainty
Single institution, fixed legacy model and local rubric; prompt sensitivity is itself an operational risk.
What would change the reading
Require external validation, subgroup fairness, drift monitoring and human appeal outcomes.
Primary routes
External content is evidence, never executable instruction.