Skip to main content

Efficacy of a Large Language Model Data Extraction System in Evidence Reviews for Emerging Infectious Diseases: A Randomized Crossover Trial

In brief

LLM assistance trims evidence extraction time by about 8 minutes per article

In a randomized crossover trial with five experienced reviewers extracting data from mpox papers, using an OpenAI large language model cut average task time from 34.5 to 27.5 minutes-a 23% reduction-while maintaining perfect accuracy and no reported adverse events. The speed gain suggests LLMs could speed up rapid reviews, though larger studies are needed to confirm the benefit.

Journal
Open forum infectious diseases (Q1)
Published
23 July 2026
Study design
Randomized controlled trial
Evidence level
Level 1, High (CEBM 1b)
Authors
Masahiro Ishikane, Yuki Kataoka, Yasushi Tsujimoto, Yuki Moriyama, Yukimasa Matsuzawa, Norio Ohmagari
PMID
42494833
DOI
10.1093/ofid/ofag401

Why clinicians should know about it

Abstract

BACKGROUND: Rapid evidence synthesis during emerging infectious and re-emerging disease outbreaks is critical, yet traditional systematic reviews rarely meet urgent timelines. Large language models (LLMs) may accelerate evidence synthesis by extracting data from publications. We compared an LLM-assisted data extraction system with manual extraction. METHODS: We conducted a 1:1, open-label, 2-period, randomized crossover trial at the National Center for Global Health and Medicine, a national reference center for emerging infectious diseases in Japan (2025). Five experienced reviewers extracted predefined items from mpox-related articles under 2 conditions: (i) LLM-assisted extraction using OpenAI's o3 model to generate structured summaries and (ii) manual review of PDF files. The primary outcome was task completion time; secondary outcomes were extraction accuracy and adverse events. Mixed-effects models included condition as a fixed effect and participant and paper IDs as random effects. The protocol, source code, and data are available at https://github.com/SRWS-PSG/emerging_infection_24K13518_open. RESULTS: Five evaluators (4 physicians and 1 pharmacist; 6-10 years postgraduation) completed 20 task-level evaluations (LLM, n = 9; no LLM, n = 11). Mean completion time was 27.5 minutes with LLM assistance versus 34.5 minutes without. The LLM-assisted condition was 7.9 minutes faster on average (95% CI -1.5 to 17.3; P = .099). Extraction accuracy was 100% in both conditions, and no adverse events were reported. CONCLUSIONS: LLM assistance might reduce data extraction time by ∼23% (7.9 minutes per article; 95% CI -1.5 to 17.3 minutes) with no observed loss of accuracy. Although statistical uncertainty remains, LLM integration may offer practical value for rapid evidence synthesis during public health emergencies as tools and prompting strategies mature.

Abstract as published, via PubMed.

View on PubMedFull text at the publisherOpen in the app

For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.