- LLM-based clustering of spontaneous speech produced clusters significantly associated with validated clinical scales and clinician diagnoses across French, Italian, and Chinese cohorts.
- The LLMs generated fine-grained, human-readable cluster descriptions, automating thematic analysis in minutes and recovering core psychiatric constructs across languages.
- Cluster associations depended on interview question; feelings-and-sleep responses showed strongest clinical links, while age and sex also influenced topic membership independently.
JMIR Ment Health. 2026 Aug 17;13:e99185. doi: 10.2196/99185.
ABSTRACT
BACKGROUND: Depression is underdiagnosed worldwide, and clinicians rely on interpreting patients’ subjective speech. Qualitative analysis of patient language does not scale, and existing computational approaches describe topics with keyword lists that miss clinical nuance.
OBJECTIVE: We evaluated whether clustering spontaneous speech transcripts with large language models (LLMs) yields clusters whose membership is associated with validated clinical scales across multilingual cohorts, and whether LLMs can render those clusters human-readable through fine-grained natural-language descriptions. We further examined which interview questions yield clusters most strongly associated with clinical status, and how sociodemographic factors relate to cluster membership.
METHODS: We analyzed spontaneous speech transcripts from 4 independent cohorts: a French general population sample (1338 participants) and 3 clinical samples in Italian (n=116), Chinese (n=52), and Spanish (n=90). Responses to open-ended questions were transcribed, embedded with a multilingual language model, dimensionally reduced, and grouped by density-based clustering. An LLM then summarized each cluster into a natural-language description. Cluster membership was tested for association with validated clinical scales (Patient Health Questionnaire-9, Beck Depression Inventory, Generalized Anxiety Disorder 7-item scale, Athens Insomnia Scale, Multidimensional Fatigue Inventory, and Columbia Suicide Severity Rating Scale), clinician-assigned depression diagnoses, and sociodemographic factors (age, education, and sex).
RESULTS: Unsupervised clustering yielded clusters significantly associated with clinical scores in the French, Italian, and Chinese cohorts, with an exploratory association in the smaller Spanish cohort. In the French general population, Patient Health Questionnaire-9 depression scores differed across clusters (η²=0.19, 95% CI 0.17 to 0.24, P<.001), as did anxiety (Generalized Anxiety Disorder 7-item scale), insomnia (Athens Insomnia Scale), and fatigue (Multidimensional Fatigue Inventory) scores (η²=0.14 to 0.16, all P<.001). In the clinical cohorts, cluster membership was associated with clinician-diagnosed depression in the Italian sample (Cramér V=0.74, 0.65 to 0.85, P<.001) and with major depressive disorder diagnosis in the Chinese sample (V=0.55, 0.36 to 0.78, P=.001). A suicide-risk association in the smaller Spanish sample (Columbia Suicide Severity Rating Scale) was unstable and is reported as exploratory (mean V=0.27). The feelings-and-sleep question yielded the strongest clinical associations, whereas questions about past or future events yielded small effect sizes (η²≤0.05). Age (η²=0.27, 0.24 to 0.32) and sex (Cramér V=0.22, 0.21 to 0.30) were each associated with cluster membership for the “describe your last 24 hours” question (both P<.001).
CONCLUSIONS: Fine-grained LLM-based topic modeling automates the labor-intensive early stages of thematic analysis, in minutes of compute, and recovers core psychiatric constructs across languages. It produces human-readable cluster descriptions. Certain interview questions yield clusters more strongly associated with clinical status than others, and sociodemographic factors shape topic content independently of clinical status, an often-overlooked confound. As the cohorts differ, cross-cohort observations are preliminary without generalizability. Screening, treatment-planning, and decision-support uses remain to be established prospectively.
PMID:42607290 | DOI:10.2196/99185
Share Evidence Blueprint
Save to Google Notes

Search Google Scholar
Save as PDF
⭐ My Revision List

