Skip to main content

A Comparison of Machine Learning and Human Graders for Glaucoma Diagnosis from Fundus Images for Population Screening

Journal
Ophthalmology (Q1)
Published
10 August 2026
Study design
Case-control study
Evidence level
Level 3, Low (CEBM 3b)
Authors
Thomas R P Taylor, Justin Khasentino, Robert N Luben, Laura Meliante, Kelsey V Stuart, Dun Jack Fu, et al.
PMID
42575300
DOI
10.1016/j.ophtha.2026.07.043

Why clinicians should know about it

  • Picked for Ophthalmology (paper of the day, 15 August 2026): ML vs human glaucoma diagnosis from fundus images

Abstract

PURPOSE: To compare the accuracy of vertical cup-disc ratios (VCDR), ascertained by machine learning (ML) versus human graders, from fundus images for glaucoma detection. This study utilizes population-based data, with a disease prevalence and case-mix that is closer to a real-world setting than conventional case-control studies, with the aim of developing improved glaucoma screening tests. DESIGN: Cross-sectional analysis of a population-based study. PARTICIPANTS: 6,304 participants of the EPIC-Norfolk Eye Study with color fundus images gradable by humans and ML in both eyes. METHODS: VCDR was independently estimated from two-dimensional fundus images of EPIC-Norfolk Eye Study participants by trained human graders (H-VCDR) and an externally trained, open access, ML model (ML-VCDR). A neural network trained on 81,830 ophthalmologist-labeled images was used to generate pseudo-labels for over 100,000 UK Biobank images, on which ML-VCDR was subsequently trained. Glaucoma status was ascertained by tertiary center specialist examination. Predictive performance of VCDR for glaucoma status was examined using logistic regression. ML-VCDR estimates were additionally compared to a popular open-source ML model (AutoMorph) and scanning laser ophthalmoscopy (Heidelberg Retinal Tomography (HRT)). MAIN OUTCOME MEASURES: Area Under the Receiver Operated Characteristic Curve (AUROC), explained variance (McFadden's pseudo-R2). RESULTS: Of 6,304 participants (mean age 68 years; 57% women), 696 had glaucoma or suspect status in at least one eye. For left eyes, H-VCDR and ML-VCDR explained 17% (95% CI 14.7 - 20.4) and 31% (95% CI 27.9 - 33.9) of glaucoma status variance and had an area under the ROC curve (AUROC) of 79% (95% CI 76.8 - 81.2) and 88% (95% CI 86.3 - 89.1), respectively. Right eye H-VCDR and ML-VCDR explained 20% (95% CI 16.9 - 22.6) and 35% (95% CI 32.4 - 37.9) of the variance, and had an AUROC of 81% (95% CI 78.6 - 82.5) and 90% (95% CI 88.7 - 91.0), respectively. ML-VCDR also performed better than AutoMorph and HRT at predicting glaucoma status from VCDR estimations. CONCLUSIONS: In this population-based setting, ML far outperformed trained human graders at predicting specialist-ascertained glaucoma status from fundus images. This provides promise for ML-supported strategies for glaucoma population screening.

Abstract as published, via PubMed.

View on PubMedFull text at the publisherOpen in the app

For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.