Diagnostic performance of AI-powered prostate MRI against biopsy ground truth: quasi-continuous risk scoring versus fixed PI-RADS thresholds
In brief
AI prostate MRI score keeps 95% sensitivity and raises specificity to 71%
In 122 patients, the AI algorithm's continuous Level-of-Suspicion score identified clinically significant prostate cancer with the same 94.7% sensitivity as the standard PI-RADS at least 4 threshold, but improved specificity from 63% to about 71%. Calibration was good and internal validation modestly reduced performance, but the difference did not reach statistical significance, so larger prospective studies are needed before routine adoption.
- Journal
- Cancer imaging : the official publication of the International Cancer Imaging Society (Q1)
- Published
- 15 September 2026
- Study design
- Prospective / inception cohort
- Evidence level
- Level 2, Moderate (CEBM 2b)
- Authors
- Nadine Bayerl, Joel Appel, Maximilian Schmidt, Alexander Cavallaro, Robert Grimm, Heinrich von Busch, et al.
- PMID
- 42745323
- DOI
- 10.1186/s40644-026-01126-5
Why clinicians should know about it
- Picked for Radiology, Radiation Oncology, Nuclear Medicine, Medical Physics and Imaging (paper of the day, 19 September 2026): AI prostate MRI diagnostic accuracy vs biopsy
Abstract
BACKGROUND: Multiparametric prostate MRI is central to prostate cancer (PCa) diagnostics, yet PI-RADS interpretation suffers from inter-reader variability. This study aimed to evaluate a commercial AI algorithm for cancer detection in prostate MRI against histopathology, comparing its standard PI-RADS classification with its quasi-continuous Level-of-Suspicion (LoS) score as decision variables. METHODS: In this retrospective single-center study, 122 mpMRI examinations, each followed by systematic and/or fusion-targeted biopsy, were analyzed. Histopathology served as the reference standard for any PCa (Gleason ≥ 6) and clinically significant PCa (csPCa; Gleason ≥ 7). Diagnostic performance of the algorithm's PI-RADS and LoS outputs was assessed by ROC analysis. AUCs were compared using DeLong's test. The Youden-optimal LoS cutoff was determined, its optimism was quantified by bootstrap internal validation, and the calibration of the LoS score was assessed after logistic recalibration. RESULTS: For csPCa, PI-RADS yielded an AUC of 0.833 (95% CI: 0.764-0.901), versus 0.882 (95% CI: 0.820-0.943) for the LoS score (p = 0.106). At PI-RADS ≥ 4, sensitivity was 94.7% and specificity 63.1%. The Youden-optimal LoS cutoff was 81. At this exploratory threshold, sensitivity remained 94.7% while specificity was 70.8%, with positive and negative predictive values of 74.0% and 93.9%, respectively. Bootstrap internal validation indicated modest optimism, with corrected estimates of 92.9% for sensitivity and 69.1% for specificity. After logistic recalibration, the bootstrap-corrected calibration slope was 0.98 with an intercept of -0.01. CONCLUSION: In this retrospective single-center cohort, the AI algorithm achieved high diagnostic accuracy for csPCa. Its quasi-continuous LoS score provided an operating point combining high sensitivity with numerically higher specificity than the PI-RADS ≥ 4 threshold. Neither this difference (exact McNemar p = 0.063) nor the difference in AUC between the two classifiers (p = 0.106) reached statistical significance, so the finding should be regarded as a promising trend rather than a demonstrated advantage. Prospective confirmation is required before AI-assisted standardized interpretation is integrated into prostate MRI workflows on this basis.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.