- GPT-4 exhibits strong convergent validity with standard instruments and expert ratings of depression (r = .70-.81).
- The models encode depression as an interconnected symptom network aligning with clinical literature, but underemphasise suicidality and overemphasise psychomotor symptoms.
- Associations imply testable hypotheses: sleep and fatigue are broadly influenced by other symptoms, whereas worthlessness and guilt link selectively to depressed mood.
J Psychopathol Clin Sci. 2026 Jul 27. doi: 10.1037/abn0001144. Online ahead of print.
ABSTRACT
The use of large language models (LLMs) such as ChatGPT (GPT-4/GPT-5) for mental health is growing rapidly, including their use to assess and help people with mood disorders, such as depression. However, we have a limited understanding of the schema of mental disorders in these language models, that is, how they internally associate and interpret the symptoms. In this work, we used contemporary measurement theory to decode how GPT-4 and GPT-5 interrelate depressive symptoms, providing an explanation of how large language models behave in clinical applications. First, we evaluated whether these models can reliably detect depressive constructs, finding that GPT-4 (a) demonstrated strong convergent validity with standard instruments and expert judgments (r = .70-.81). We mapped the model’s internal schema of depression, showing that it (b) organized symptoms into an interconnected network (symptom correlations r = .23-.78) that largely aligned with the established clinical literature. However, it (c) underemphasized the relationship between suicidality and other symptoms while overemphasizing the importance of psychomotor symptoms. Finally, the associations (d) suggested novel hypotheses for symptom mechanisms, for instance, indicating that sleep and fatigue are broadly influenced by other depressive symptoms, whereas worthlessness/guilt is selectively associated with depressed mood. GPT-5 showed slightly lower convergence with self-report, which our analysis attributed to changes in the relationships among symptoms. These insights provide a generalizable, empirical foundation for understanding language models’ mental health assessments across models and disorders. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
PMID:42507384 | DOI:10.1037/abn0001144
Share Evidence Blueprint
Save to Google Notes

Search Google Scholar
Save as PDF
⭐ My Revision List

