Benson Mwangi, Hammza Jabbar Abdl Sattar Hamoudi, Marsal Sanches, Nurhak Doğan, Pooja Chaudhary, Mon-Ju Wu, Giovana B. Zunta-Soares, Jair C. Soares, Andrés Martin, César A. Soutullo · npj Mental Health Research 2026 · 2026
DOI: 10.1038/s44184-026-00245-y
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The Mental Status Examination (MSE) relies on complex clinical reasoning, presenting a unique challenge for artificial intelligence (AI) benchmarking. We introduce a multicenter framework evaluating an open-weight multimodal foundation model (Qwen3-Omni) against independent expert panels at UTHealth and Yale using simulated clinical videos. Across 396 classifications spanning 10 MSE domains, inter-expert agreement was substantial (Gwet’s AC1 = 0.87), whereas model-expert alignment was moderate (AC1 = 0.70-0.72). Crucially, a structured reasoning gap emerged where the model over-predicted pathology in observable domains (e.g., speech, affect) but systematically under-detected findings requiring inferential interpretation of subjective experiences (e.g., delusions). Additionally, model agreement declined as symptom severity increased, whereas inter-expert agreement remained substantially higher. Parameter quantization degraded performance heterogeneously across domains, while an architecturally distinct model (Phi-4-Multimodal-Instruct) exhibited similar domain-specific errors. These findings support the use of inter-expert agreement as a calibration benchmark for psychiatric AI. They reveal that the primary barrier to clinical translation is no longer perceptual feature extraction, but the higher-order inferential reasoning required to map patient narratives onto psychiatric constructs.
No comments yet — start the discussion below.