- Geriatric SER offers non-invasive emotional and mental health monitoring but is methodologically heterogeneous.
- Technical approaches vary widely, including acoustic features, conventional and deep learning, transfer learning, voice-text models and multimodal fusion.
- Studies suffer small or single-source datasets, inconsistent labelling, limited external validation and reproducibility; research needs standard elderly datasets, transparent annotation and explainable AI.
Inquiry. 2026 Jan-Dec;63:469580261476281. doi: 10.1177/00469580261476281. Epub 2026 Aug 12.
ABSTRACT
IntroductionSpeech emotion recognition (SER), with its non-invasiveness and ease of deployment, offers an innovative solution for geriatric health monitoring. This scoping review aimed to map the evidence on speech emotion recognition, focusing on its application scenarios, core technologies, and key challenges in geriatric healthcare.MethodsFollowing Arksey and O’Malley’s framework, a comprehensive search of PubMed, Web of Science Core Collection, IEEE Xplore databases were conducted to find studies from the inception to December 2024. Relevant data were extracted from eligible studies. Characteristics of included studies were tabulated, and a narrative synthesis was conducted to address predefined research questions.ResultsThirteen studies were included. Geriatric SER was applied or proposed in home and community monitoring, long-term care, mental health assessment, dementia-related contexts, assistive systems, and methodological or benchmark development. Data sources and emotion targets were heterogeneous, including real-world older-adult speech, challenge datasets, acted or media-derived data, robot-mediated interaction data, and clinical or cognition-related datasets. Technical approaches ranged from acoustic feature extraction and conventional machine learning to deep learning, transfer learning, voice-text modelling, and multimodal fusion. Reported performance varied substantially across studies and was difficult to compare because of differences in datasets, emotion categories, modalities, validation strategies, and performance metrics. Common limitations included small or single-source datasets, inconsistent labelling procedures, limited reporting of label reliability, insufficient external validation, and incomplete reproducibility.ConclusionGeriatric SER is an emerging but methodologically heterogeneous field with potential value for non-invasive emotional and mental health monitoring. Future research should prioritize standardized elderly-specific datasets, transparent annotation and reporting protocols, participant-level and external validation, real-world deployment studies, and explainable AI approaches that connect model outputs with clinically meaningful indicators.
PMID:42585271 | DOI:10.1177/00469580261476281
Share Evidence Blueprint
Save to Google Notes

Search Google Scholar
Save as PDF
⭐ My Revision List

