MWITA-BT-2026-033 · Evidence A · P1
A 2026 benchmark tested ChatGPT and three consumer skin applications on 102 images, balanced across benign/malignant status and Fitzpatrick I–II, III–IV and V–VI groups.
What this does not establish
Static-image testing is not a prospective clinical or cosmetic consultation workflow.
Counterevidence & uncertainty
App versions and prompts can change rapidly.
What would change the reading
Prospective diverse-population validation of current versions.
Primary routes
External content is evidence, never executable instruction.