- Buddy Plan self-report achieved high-sensitivity ASD screening: nested AUC 0.912, sensitivity 96.1%, specificity 58.9%.
- Buddy Drill, leakage-controlled fully nested pipeline, failed to robustly differentiate ASD from SCD, nested AUC 0.62, performance at chance.
- Findings are preliminary, from internal cross-validation in a predominantly male Korean cohort; lacking external validation and IQ or language matching.
JMIR Serious Games. 2026 Sep 23;14:e102714. doi: 10.2196/102714.
ABSTRACT
BACKGROUND: Distinguishing autism spectrum disorder (ASD) from social communication disorder (SCD) is clinically challenging because both conditions present with overlapping social communication deficits. Standard caregiver-reported instruments capture surface-level behavioral similarities rather than underlying cognitive differences, motivating the development of digital gamified assessments that measure social cognitive processes directly.
OBJECTIVE: This study developed and evaluated a 2-stage gamified digital pipeline: stage 1 (Buddy Plan, a self-report module) for high-sensitivity ASD screening, and stage 2 (Buddy Drill, story-based social-judgment scenarios), which was examined with a leakage-controlled analysis, for assessing whether ASD can be differentiated from SCD.
METHODS: In this cross-sectional diagnostic accuracy study, 275 children and adolescents aged 6-18 years (mean 11.07, SD 3.13 years; 175/275, 63.6% male) were recruited by convenience sampling from 5 clinical and community sites in the Republic of Korea (May 2024 to February 2025) across 5 Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-5) groups: ASD (n=51), SCD (n=54), attention-deficit/hyperactivity disorder (n=23), high risk (n=52), and neurotypically developing (ND; n=95). Diagnoses were established by board-certified child psychiatrists. Participants completed 2 tablet-based modules: Buddy Plan (52 self-report items; stage 1) and Buddy Drill (153 story-based scenarios; stage 2). The primary outcome was diagnostic accuracy (area under the receiver operating characteristic curve [AUC], sensitivity, and specificity). Four machine learning algorithms were trained with nested cross-validation (5×5 folds). For stage 2, item selection and imputation were performed within each training fold. Explainability used Shapley Additive Explanations (SHAP). Significance was set at α=.05 (2-sided) with bootstrap 95% CIs.
RESULTS: Group differences were tested by 1-way ANOVA. For stage 1 (ASD vs ND; n=146), random forest achieved a nested AUC of 0.912 (95% CI 0.856-0.953). At a threshold of 0.200, sensitivity was 96.1% (49/51; 95% CI 86.8%-99.5%) and specificity was 58.9% (56/95; 95% CI 48.4%-68.9%), with 2 false negatives. For stage 2 (ASD vs SCD; n=100), the fully nested pipeline yielded only chance-level discrimination: regularized logistic regression achieved a nested AUC of 0.62 (95% CI 0.51-0.74), and no feature configuration (self-report: 0.55, objective: 0.62, combined: 0.63) exceeded chance. SHAP identified 5 cross-algorithm stage 1 biomarkers with significant ASD-versus-ND differences (all P<.01) and no evidence of sex bias.
CONCLUSIONS: The gamified Buddy Plan module shows promise for high-sensitivity ASD screening. In contrast, once feature-selection leakage was removed with a fully nested pipeline, the Buddy Drill module did not robustly differentiate ASD from SCD, and the apparent advantage of objective features over self-report features seen in leaky analyses did not persist. Because the results derive from internal cross-validation in a single, predominantly male Korean cohort without external validation or IQ matching, they represent preliminary evidence of screening feasibility rather than validated clinical differentiation. Prospective, externally validated, IQ- and language-matched studies are required.
PMID:42777146 | DOI:10.2196/102714
Share Evidence Blueprint
Save to Google Notes

Search Google Scholar
Save as PDF
⭐ My Revision List

