- DSM-5 checklist model achieved high internal agreement: 92.5% accuracy and ROC-AUC 0.984, indicating strong internal performance but not clinical validity.
- Facial-image AI performance varied markedly by model, training data, and population transfer; local adaptation improved accuracy but results depended on architecture and small test cohorts.
- Findings are preliminary; AI cannot replace comprehensive clinical assessment and require prospective multicentre, participant-level evaluation with independent reference standards before clinical implementation.
Health Sci Rep. 2026 Oct 8;9(10):e73360. doi: 10.1002/hsr2.73360. eCollection 2026 Oct.
ABSTRACT
BACKGROUND AND AIMS: Artificial intelligence methods may support research on autism spectrum disorder (ASD)-related screening, but they cannot replace comprehensive clinical assessment. We evaluated two image-training scenarios public Kaggle data alone and a manually prepared combined training set containing Kaggle images plus local images and separately examined a structured DSM-5 checklist model. A child-level reanalysis was used to evaluate public-to-local transfer and local adaptation while ensuring that images from the same child were not assigned to different partitions.
METHODS: This methodological study used 268 retrospectively abstracted DSM-5 records (134 ASD and 134 non-ASD) and 162 local facial images from 92 children (43 ASD and 49 non-ASD). The image experiment evaluated EfficientNet-B2 and ResNet50 under two training scenarios: Kaggle-only training with 2536 images (1268 per class) and a manually prepared combined training set with 2580 images (1290 per class), formed by adding 22 local images per class to the Kaggle training folders. Local images were linked to anonymous child identifiers before partitioning.
RESULTS: The checklist model achieved 92.5% accuracy (248/268; 95% CI, 89.6%-95.5%) and ROC-AUC 0.984 (95% CI, 0.971-0.993) in internal out-of-fold evaluation. In the child-level reanalysis, the Kaggle-only scenario produced zero-shot/adapted accuracies of 42.1%/68.4% for EfficientNet-B2% and 68.4%/89.5% for ResNet50 in 19 locked test children. With the 2580-image combined training set, zero-shot/adapted accuracies were 71.4%/78.6% for EfficientNet-B2% and 57.1%/85.7% for ResNet50 in 14 locked test children. These image results are preliminary and based on small local test cohorts.
CONCLUSION: The checklist analysis showed high internal agreement with an existing psychiatrist-assigned label, whereas image performance varied by architecture, training scenario, population transfer, and local adaptation. The results do not establish clinical validity or autonomous diagnosis. Future prospective, multicenter, participant-level evaluation with independent reference standards is required before any clinical use.
PMID:42857342 | PMC:PMC13650386 | DOI:10.1002/hsr2.73360
Share Evidence Blueprint
Save to Google Notes

Search Google Scholar
Save as PDF
⭐ My Revision List

