- Large language models can detect changes in depression severity from digital phenotyping data and outperform traditional baseline models.
- Both few-shot prompted and fine-tuned LLMs perform well; embedding-only and QLoRA fine-tuning show comparable results with differing strengths.
- LLMs offer promise for integrating heterogeneous behavioural data, but require clinical validation and ethical safeguards for real-world deployment.
NPJ Digit Med. 2026 Jun 15. doi: 10.1038/s41746-026-02883-0. Online ahead of print.
ABSTRACT
Digital phenotyping leverages continuous data from smartphones and wearable devices for real-time mental health monitoring, offering opportunities for early detection and personalized care in mood disorders by enabling clinicians to proactively respond to significant changes before symptoms worsen. However, the heterogeneity of such data presents modeling challenges. This study evaluates the potential of large language models (LLMs) to detect changes in depression severity from digital phenotyping data among individuals experiencing major depressive episodes. We compare in-context learning and fine-tuning strategies and find that both few-shot prompted and fine-tuned LLMs outperform traditional baselines. Furthermore, embedding-only and QLoRA fine-tuning yield comparable results, with the former excelling on individual features and the latter performing better on combined inputs. These findings demonstrate the promise of LLMs in integrating heterogeneous behavioral data for mental health analysis, while underscoring the importance of clinical validation and ethical safeguards in real-world applications.
PMID:42298144 | DOI:10.1038/s41746-026-02883-0
Share Evidence Blueprint
Save to Google Notes

Search Google Scholar
Save as PDF

