- Proposed a scalable, agnostic visual framework to assess fidelity of healthcare relational databases by systematically analysing multivariate associations among coded categorical variables.
- Converted categorical variables to binary features, computed thousands of pairwise association metrics, and summarised results with curves, bubble charts, heatmaps and quantitative coefficients.
- Framework detected gradual fidelity loss under simulated degradations, offering both general overview and granular insights suitable for evaluating synthetic databases.
Int J Med Inform. 2026 Sep 3;222:106711. doi: 10.1016/j.ijmedinf.2026.106711. Online ahead of print.
ABSTRACT
BACKGROUND: Although data reuse is increasingly common in healthcare research, changing regulatory frameworks are impeding these efforts. Synthetic data constitute a promising issue but also create challenges, such as the assessment of data fidelity and the ability to output identical statistical results. Conventional validation approaches often rely on univariate or bivariate comparisons, which fail to capture the complexity of associations between multivalued categorical variables.
MATERIAL: We built and studied a fictitious database of 10,000 hospital stays reproducing the structure of the French Programme de Médicalisation des Systèmes d’Information database. Each stay included single-valued variables (one value per individual: sex, age in deciles, and diagnosis-related group) and multivalued variables (zero, one or several values per individual: diagnoses coded according to the International Classification of Diseases, 10th Edition, and procedures coded according to the French Classification Commune des Actes Médicaux).
METHOD: All categorical variables were binarized, and thousands of pairwise association metrics (primarily odds ratios) were calculated for the reference and evaluation datasets. The results were summarized using curves, bubble charts, heatmaps, and coefficients such as exponential mean deviation. Simulated data degradations from 0% to 100% were introduced to evaluate the method’s sensitivity.
RESULTS: We analyzed 500 ICD-10 diagnoses and 450 CCAM procedures, representing 225,000 possible combinations. We developed and evaluated graphical representations for assessing data fidelity at a glance. In simulations of an increasing degree of data degradation, those graphical representations and comprehensive, quantitative metrics facilitated the detection of the gradual loss of data fidelity.
CONCLUSION: We developed a simple, scalable, agnostic framework for assessing the fidelity of healthcare databases by systematically analyzing associations among the modalities of coded variables. This method complements existing approaches. It is particularly suitable for the evaluation of synthetic relational databases because it offers both general and granular insights into data fidelity loss.
PMID:42700765 | DOI:10.1016/j.ijmedinf.2026.106711
Share Evidence Blueprint
Save to Google Notes

Search Google Scholar
Save as PDF
⭐ My Revision List

