Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases
In brief
Cornea-focused AI raises diagnostic success to roughly three-quarters of cases
In a masked trial of 39 real-world corneal cases, a retrieval-augmented, cornea-specialized GPT-4o model lifted clinicians' diagnostic accuracy from about 49% unaided to 72% overall, outperforming the standard GPT-4o. The boost was strongest for trainees with lower baseline performance, while its effect on management decisions was mixed, highlighting the need for further testing of AI-assisted treatment planning.
- Journal
- International ophthalmology (Q2)
- Published
- 2 September 2026
- Study design
- Randomized controlled trial
- Evidence level
- Level 1, High (CEBM 1b)
- Authors
- David Mikhail, Daniel Milad, Fares Antaki, Jason Milad, Fady Sedarous, Rachel Ann Martin, et al.
- PMID
- 42684491
- DOI
- 10.1007/s10792-026-04232-2
Why clinicians should know about it
- Picked for Ophthalmology (top studies of the week, 6 September 2026): LLM improves diagnostic accuracy in complex corneal cases
Abstract
PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.