MWITA-SR-2026-012 · Evidence A · P0
Fine-tuning on response distributions outperformed tested prompting and zero-shot baselines on seen and unseen benchmark splits.
What this does not establish
Relative improvement does not establish sufficient validity for decisions.
What would change the reading
Recalibrate on held-out, recent human data after every material model change.
Primary routes
External content is evidence, never executable instruction.