Skip to main content

External Validation of Clinical Risk Scores and Machine Learning Models for Predicting 30-Day Cardiovascular Risk After Noncardiac Surgery: The PERICARE Study

Journal
Journal of clinical medicine (Q1)
Published
4 September 2026
Study design
Prospective / inception cohort
Evidence level
Level 2, Moderate (CEBM 2b)
Authors
Aslan Erdoğan, Şeyma Yeşil, Gamze Gençol Akçay, Ufuk Sali Halil, İhsan Demirtaş, Ezgi Alp, et al.
PMID
42739853
DOI
10.3390/jcm15176848

Why clinicians should know about it

  • Picked for Health Informatics (paper of the day, 16 September 2026): External validation of ML risk scores

Abstract

Background: Machine learning (ML) has emerged as a promising approach for preoperative cardiovascular risk prediction; however, the generalizability of ML models across institutions remains uncertain. Moreover, comprehensive head-to-head comparisons between ML algorithms and established clinical risk scores for predicting 30-day major adverse cardiac events (MACE) are scarce. We therefore evaluated the performance and external transportability of multiple ML models across independent centers and compared their predictive accuracy with validated benchmark clinical risk scores. Methods: In a site-separated two-center cohort (derivation n = 707, 27 MACE; external validation n = 378, 38 MACE), ten algorithms trained on preoperative variables were externally validated without refitting and benchmarked against the American university of Beirut-HAS2 (AUB-HAS2), American society of anesthesiologists (ASA), and revised cardiac risk index (RCRI). We assessed AUROC, calibration, Brier score, and decision-curve net benefit, with paired bootstrap comparisons, DeLong testing, IDI, and NRI. Results: Among the ML models, no single algorithm consistently outperformed the others across all performance metrics. Naive Bayes achieved the highest external discrimination (AUROC 0.738, 95% CI 0.668-0.804) but showed poor calibration and threshold-dependent clinical utility. Gradient Boosting showed the most favorable balance of discrimination and calibration slope (AUROC 0.707; calibration slope 0.991), although absolute risk remained underestimated in external validation, whereas HistGradient Boosting yielded the best overall probability prediction, with the lowest Brier score (0.087) and the greatest decision-curve net benefit at clinically relevant risk thresholds. However, in external validation, no ML model demonstrated statistically significant AUROC superiority over the AUB-HAS2 score (all p > 0.05), although several models modestly outperformed the RCRI. Conclusions: In this site-separated external validation, ML models showed metric-dependent performance but no discrimination advantage over the AUB-HAS2 index. Given low event counts, flexible-model results are hypothesis-generating. These findings provide a cautionary, reproducible benchmark; local recalibration and prospective evaluation are prerequisites before clinical deployment.

Abstract as published, via PubMed.

View on PubMedFull text at the publisherOpen in the app

For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.