Reducing Clinician Annotation Fatigue in Open-Set Federated Learning: A Client-Adaptive Vision-Language Gatekeeper
In brief
A mammography AI gate cut annotation time by an estimated 20% after false removals
In a retrospective, artifact-enriched cohort, a scanner-adaptive AI gate was estimated to reduce annotation time by 20% when only correctly removed artifacts counted; it withheld about two-thirds of artifacts but also removed 15.6% of relevant images. The study involved one radiologist, and real-world time savings may be smaller, so multireader testing is needed before deployment.
- Journal
- Journal of the American College of Radiology : JACR (Q1)
- Published
- 1 October 2026
- Study design
- Prospective / inception cohort
- Evidence level
- Level 2, Moderate (CEBM 2b)
- Authors
- Adea Nesturi, David Dueñas Gaviria, Jiajun Zeng, Mohammed Bahaaeldin, Julian A Luetkens, Shadi Albarqouni
- PMID
- 42822685
- DOI
- 10.1016/j.jacr.2026.09.023
Why clinicians should know about it
- Picked for Radiology, Radiation Oncology, Nuclear Medicine, Medical Physics and Imaging (paper of the day, 4 October 2026): Federated learning gatekeeper to reduce annotation fatigue in breast imaging
Abstract
PURPOSE: Federated learning enables breast-imaging sites to jointly train mammography artificial intelligence (AI) without sharing images, but radiologists at each site must still annotate selected images during active-learning rounds. We developed and evaluated a client-adaptive vision-language gatekeeper that withholds radiopaque artifact-containing images considered unsuitable for a breast-density annotation objective and quantified its effect on annotation time, queue quality, and label reliability. METHODS: In this retrospective, scanner-partitioned study, we used FedEMBED, derived from the public Emory Breast Imaging Dataset (222,700 training and 65,891 test images; four scanners; artifact prevalence 4.8%-13.1% per scanner). PromptGate adapted a biomedical vision-language model with site-specific learnable prompts and classified each image as relevant in-distribution (ID) or irrelevant out-of-distribution (OOD) using the class with the highest softmax probability. A single radiologist who is a member of the study team and has 2 years of breast imaging experience independently read a patient-separated, artifact-enriched cohort of 500 single mammographic views (ID, n=346; OOD, n=154). Cases were randomized, ID and OOD images were intermixed, and the reader was blinded to reference labels, PromptGate decisions, enrichment status, and scanner metadata. RESULTS: PromptGate withheld 101 of 154 artifact-containing images and 54 of 346 ID images, corresponding to 65.6% artifact sensitivity, 84.4% ID retention, and a 15.6% false-removal rate among ID images. Gross annotation time decreased by ∼30% (3.82 to 2.68 hours; 1.14 hours) in the enriched cohort; an analysis crediting only correctly removed artifacts yielded an estimated 20.0% reduction. A simple prevalence-based extrapolation suggests an approximate 20- to 65-min saving per 500 images at the observed scanner-specific prevalence, although the absolute benefit is expected to be smaller than in the enriched cohort. Artifact discrimination reached an area under the receiver operating characteristic curve of 0.85. Reader agreement with the manual artifact reference was 93.6% (κ=0.85), whereas agreement with the clinical-record density label was 60.9% exact (κ=0.48). Queue purity reached 92.5%, compared with 88.5% without filtering and 87.0% with a static vision-language model filter, while downstream density accuracy remained similar to the unfiltered condition. CONCLUSION: In this retrospective proof-of-concept study, a privacy-oriented, scanner-adaptive vision-language gate improved the concentration of task-relevant images in federated breast-density annotation queues and reduced gross reading time. Because some ID images were incorrectly withheld and the reader study used an artifact-enriched cohort, a prospective multireader and multi-institutional evaluation is needed before operational deployment.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.