Efficacy of a Large Language Model Data Extraction System in Evidence Reviews for Emerging Infectious Diseases: A Randomized Crossover Trial
In brief
LLM assistance trims evidence extraction time by about 8 minutes per article
In a randomized crossover trial with five experienced reviewers extracting data from mpox papers, using an OpenAI large language model cut average task time from 34.5 to 27.5 minutes-a 23% reduction-while maintaining perfect accuracy and no reported adverse events. The speed gain suggests LLMs could speed up rapid reviews, though larger studies are needed to confirm the benefit.
- Journal
- Open forum infectious diseases (Q1)
- Published
- 23 July 2026
- Study design
- Randomized controlled trial
- Evidence level
- Level 1, High (CEBM 1b)
- Authors
- Masahiro Ishikane, Yuki Kataoka, Yasushi Tsujimoto, Yuki Moriyama, Yukimasa Matsuzawa, Norio Ohmagari
- PMID
- 42494833
- DOI
- 10.1093/ofid/ofag401
Why clinicians should know about it
- Picked for Infectious Diseases (paper of the day, 25 July 2026).
Abstract
BACKGROUND: Rapid evidence synthesis during emerging infectious and re-emerging disease outbreaks is critical, yet traditional systematic reviews rarely meet urgent timelines. Large language models (LLMs) may accelerate evidence synthesis by extracting data from publications. We compared an LLM-assisted data extraction system with manual extraction. METHODS: We conducted a 1:1, open-label, 2-period, randomized crossover trial at the National Center for Global Health and Medicine, a national reference center for emerging infectious diseases in Japan (2025). Five experienced reviewers extracted predefined items from mpox-related articles under 2 conditions: (i) LLM-assisted extraction using OpenAI's o3 model to generate structured summaries and (ii) manual review of PDF files. The primary outcome was task completion time; secondary outcomes were extraction accuracy and adverse events. Mixed-effects models included condition as a fixed effect and participant and paper IDs as random effects. The protocol, source code, and data are available at https://github.com/SRWS-PSG/emerging_infection_24K13518_open. RESULTS: Five evaluators (4 physicians and 1 pharmacist; 6-10 years postgraduation) completed 20 task-level evaluations (LLM, n = 9; no LLM, n = 11). Mean completion time was 27.5 minutes with LLM assistance versus 34.5 minutes without. The LLM-assisted condition was 7.9 minutes faster on average (95% CI -1.5 to 17.3; P = .099). Extraction accuracy was 100% in both conditions, and no adverse events were reported. CONCLUSIONS: LLM assistance might reduce data extraction time by ∼23% (7.9 minutes per article; 95% CI -1.5 to 17.3 minutes) with no observed loss of accuracy. Although statistical uncertainty remains, LLM integration may offer practical value for rapid evidence synthesis during public health emergencies as tools and prompting strategies mature.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.