Transforming Systematic Reviews: Evaluating a Fine-Tuned Large Language Model for Abstract Screening in Uveitis and Retinal Vasculitis: Fine-Tuned LLM for Review Screening
In brief
Fine-tuned LLM achieves 93% accuracy screening retinal vasculitis abstracts
In a head-to-head test of 1,030 papers, the domain-specific model UveAItis correctly classified 93.3% of titles and abstracts, far surpassing generic GPT-4o, Claude Sonnet and final-year medical students. The tool identified nearly two-thirds of expert-selected studies while generating far fewer "need consensus" flags, suggesting it could speed systematic reviews in ophthalmology, though real-world workflow integration remains to be proven.
- Journal
- Ophthalmology science (Q1)
- Published
- 22 June 2026
- Study design
- Systematic review of cohort studies
- Evidence level
- Level 2, Moderate (CEBM 2a)
- Authors
- Carlos Cifuentes-González, Maxwell B Singer, William Rojas-Carabali, Germán Mejía-Salgado, Maria Vittoria Cicinelli, Jyotirmay Biswas, et al.
- PMID
- 42620648
- DOI
- 10.1016/j.xops.2026.101296
Why clinicians should know about it
- Picked for Ophthalmology (top studies of the week, 23 August 2026): Fine‑tuned LLM improves systematic review screening in uveitis/retinal vasculitis
Abstract
PURPOSE: To evaluate the classification performance of UveAItis, a domain-specific large language model (LLM) fine-tuned for automated title and abstract screening in systematic reviews, using retinal vasculitis as a prototype. DESIGN: Comparative evaluation study embedded within a registered systematic review and meta-analysis (PROSPERO: CRD42023489232). SUBJECTS: A total of 1030 randomly selected articles from an initial search of 5533 records related to retinal vasculitis. METHODS: Articles were independently screened by 2 uveitis experts (gold standard), final-year medical students, and 3 LLMs: UveAItis (fine-tuned Generative Pre-trained Transformer [GPT]-4o), base GPT-4o, and Claude Sonnet 3.5. Screening followed a 2-question binary logic regarding human subjects and primary empirical research design. Discrepancies were resolved through expert adjudication. MAIN OUTCOME MEASURES: Classification accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and Cohen Kappa coefficient for inter-rater agreement. RESULTS: UveAItis achieved the highest performance with an accuracy of 93.3%, area under the curve (AUC) of 0.887, and Kappa of 0.77. It significantly outperformed base GPT-4o (AUC: 0.805, P = 0.021), Claude Sonnet 3.5 (AUC: 0.669, P < 0.0001), and medical students (AUC: 0.585, P < 0.00001). The fine-tuned model correctly identified 65.4% of expert-included articles postconsensus, whereas students only identified 22.3%. UveAItis also demonstrated the lowest rate of ambiguous "Need Consensus" outputs (3.4%) compared to experts (16.7%). CONCLUSIONS: UveAItis demonstrated expert-level performance, significantly outperforming general-purpose LLMs and nonexpert human reviewers. These findings validate the potential of domain-specific fine-tuning to enhance the efficiency, scalability, and reproducibility of evidence synthesis in specialized medical fields like ophthalmology. FINANCIAL DISCLOSURES: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.