- Combining multi-site studies creates structured missingness, where deterministic variable gaps break assumptions of standard imputation methods.
- Popular imputation algorithms systematically distort data distributions and fail under structured missingness, with conventional accuracy metrics often unable to detect these biases.
- Distribution-preserving methods and hierarchical modelling that account for site differences improve robustness; plus a multi-metric evaluation and decision guide for method selection.
Commun AI Comput. 2026;1(1):19. doi: 10.1038/s44488-026-00025-9. Epub 2026 Oct 9.
ABSTRACT
Missing data are ubiquitous in biomedical research, and become particularly problematic when combining data from multiple clinical studies, an increasingly common practice for building models that generalise across populations. This process introduces complex structured patterns of missing data (‘structured missingness’), where one study may omit variables that another collects, leaving deterministic gaps that standard imputation methods were not designed to handle. Here, we show that many popular imputation algorithms systematically fail under structured missingness, distorting the underlying statistical properties of the data in ways that conventional accuracy metrics cannot detect. Methods designed to preserve data distributions are more robust, and a hierarchical modelling approach that explicitly accounts for site-specific differences further improves performance. We introduce a multi-metric evaluation framework and a practical decision guide to support method selection. Our findings highlight the need for more principled imputation approaches in multi-site biomedical research and provide concrete tools to address this challenge.
PMID:42859516 | PMC:PMC13652732 | DOI:10.1038/s44488-026-00025-9
Share Evidence Blueprint
Save to Google Notes

Search Google Scholar
Save as PDF
⭐ My Revision List

