Machine learning-enabled prediction of ART pregnancy outcomes: a systematic review and meta-analysis
In brief
Machine-learning models predict ART pregnancy with about 75% sensitivity and 80% specificity
A meta-analysis of 14 studies found AI tools correctly identified roughly three-quarters of successful ART cycles and correctly ruled out about four-fifths of failures, showing moderate overall accuracy. However, most studies had high or unclear bias and large heterogeneity, so larger prospective, multicenter validations are needed before routine clinical use.
- Journal
- Journal of assisted reproduction and genetics (Q1)
- Published
- 6 August 2026
- Study design
- Systematic review of cohort studies
- Evidence level
- Level 2, Moderate (CEBM 2a)
- Authors
- Biying Li, Hong Liu, Fan Yu, Mai Xiong, Ting Tang, Rong-Hua Wu, et al.
- PMID
- 42560469
- DOI
- 10.1007/s10815-026-03988-x
Why clinicians should know about it
- Picked for Genetics (clinical) (top studies of the week, 9 August 2026).
- Picked for Reproductive Medicine (top studies of the week, 9 August 2026).
Abstract
OBJECTIVE: To systematically evaluate the diagnostic accuracy and methodological quality of machine learning (ML) prediction models for pregnancy outcomes after assisted reproductive technology (ART). METHODS: PubMed, Embase, the Cochrane Library, IEEE Xplore, MEDLINE, ClinicalTrials.gov, CNKI, Wanfang, and VIP were searched from inception to July 2026. Eligible studies developed or validated ML models to predict clinical pregnancy or live birth after ART. For studies reporting complete 2 × 2 contingency data, pooled sensitivity, specificity, diagnostic odds ratio (DOR), and summary receiver operating characteristic (SROC) curves were estimated using random-effects diagnostic meta-analysis. Risk of bias was assessed with PROBAST. RESULTS: Twenty studies were included in the systematic review, of which 14 contributed to the diagnostic meta-analysis. Overall risk of bias was low in 1 study (5.0%), high in 8 studies (40.0%), and unclear in 11 studies (55.0%). The pooled sensitivity was 0.737 (95% CI, 0.662-0.799) and the pooled specificity was 0.789 (95% CI, 0.709-0.851), with substantial heterogeneity (I2 = 97.7% and 99.0%, respectively). The pooled DOR was 10.49 (95% CI, 6.28-17.53), and the SROC curve indicated acceptable overall discrimination. Exploratory DOR subgroup analyses showed comparable performance for clinical pregnancy and live birth. No statistically robust subgroup difference was observed by algorithm type, center type, or validation status under a random-effects framework; study design showed a significant subgroup difference, but this estimate was driven by a single prospective study. CONCLUSION: ML models show moderate diagnostic accuracy for predicting ART pregnancy outcomes, but the evidence base is limited by substantial heterogeneity and frequent high or unclear risk of bias. Future studies should follow TRIPOD + AI and PROBAST-aligned standards, report calibration and clinical utility, and prioritize prospective multi-center external validation before clinical implementation. SYSTEMATIC REVIEW REGISTRATION: PROSPERO, CRD420251108846.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.