Skip to main content

Machine learning-enabled prediction of ART pregnancy outcomes: a systematic review and meta-analysis

In brief

Machine-learning models predict ART pregnancy with about 75% sensitivity and 80% specificity

A meta-analysis of 14 studies found AI tools correctly identified roughly three-quarters of successful ART cycles and correctly ruled out about four-fifths of failures, showing moderate overall accuracy. However, most studies had high or unclear bias and large heterogeneity, so larger prospective, multicenter validations are needed before routine clinical use.

Journal
Journal of assisted reproduction and genetics (Q1)
Published
6 August 2026
Study design
Systematic review of cohort studies
Evidence level
Level 2, Moderate (CEBM 2a)
Authors
Biying Li, Hong Liu, Fan Yu, Mai Xiong, Ting Tang, Rong-Hua Wu, et al.
PMID
42560469
DOI
10.1007/s10815-026-03988-x

Why clinicians should know about it

Abstract

OBJECTIVE: To systematically evaluate the diagnostic accuracy and methodological quality of machine learning (ML) prediction models for pregnancy outcomes after assisted reproductive technology (ART). METHODS: PubMed, Embase, the Cochrane Library, IEEE Xplore, MEDLINE, ClinicalTrials.gov, CNKI, Wanfang, and VIP were searched from inception to July 2026. Eligible studies developed or validated ML models to predict clinical pregnancy or live birth after ART. For studies reporting complete 2 × 2 contingency data, pooled sensitivity, specificity, diagnostic odds ratio (DOR), and summary receiver operating characteristic (SROC) curves were estimated using random-effects diagnostic meta-analysis. Risk of bias was assessed with PROBAST. RESULTS: Twenty studies were included in the systematic review, of which 14 contributed to the diagnostic meta-analysis. Overall risk of bias was low in 1 study (5.0%), high in 8 studies (40.0%), and unclear in 11 studies (55.0%). The pooled sensitivity was 0.737 (95% CI, 0.662-0.799) and the pooled specificity was 0.789 (95% CI, 0.709-0.851), with substantial heterogeneity (I2 = 97.7% and 99.0%, respectively). The pooled DOR was 10.49 (95% CI, 6.28-17.53), and the SROC curve indicated acceptable overall discrimination. Exploratory DOR subgroup analyses showed comparable performance for clinical pregnancy and live birth. No statistically robust subgroup difference was observed by algorithm type, center type, or validation status under a random-effects framework; study design showed a significant subgroup difference, but this estimate was driven by a single prospective study. CONCLUSION: ML models show moderate diagnostic accuracy for predicting ART pregnancy outcomes, but the evidence base is limited by substantial heterogeneity and frequent high or unclear risk of bias. Future studies should follow TRIPOD + AI and PROBAST-aligned standards, report calibration and clinical utility, and prioritize prospective multi-center external validation before clinical implementation. SYSTEMATIC REVIEW REGISTRATION: PROSPERO, CRD420251108846.

Abstract as published, via PubMed.

View on PubMedFull text at the publisherOpen in the app

For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.