Skip to main content

A Modular Multi-Agent Reinforcement Learning Framework Guided by LLMs: Improving Quality and Efficiency of Treatment Planning in Lung Radiotherapy

Journal
International journal of radiation oncology, biology, physics (Q1)
Published
17 September 2026
Study design
Unclassified
Evidence level
Level 5, Expert Opinion (CEBM 5)
Authors
Zipai Wang, Hao Guo, Yang Lei, Robert Samstein, Kenneth E Rosenzweig, Ming Chao, et al.
PMID
42754176
DOI
10.1016/j.ijrobp.2026.09.008

Why clinicians should know about it

  • Picked for Medical Physics (paper of the day, 20 September 2026): Hybrid LLM‑RL framework improves lung IMRT planning efficiency

Abstract

PURPOSE: Knowledge-based planning (KBP) frequently requires manual refinement to satisfy institution-specific clinical constraints, while existing automated approaches based on large language models (LLMs) or reinforcement learning (RL) each face distinct limitations in long-term optimization awareness and scalability. We propose a hybrid LLM-guided modular RL framework that integrates the strengths of both approaches for automated lung IMRT treatment planning. METHODS: The framework employs a two-level hierarchical architecture in which task-specific RL agents execute dose modifications under the coordination of an LLM supervisory layer comprising Planner, Evaluator, and Supervisor agents. Modular parallel and serial OAR agents are trained using SARSA with linear function approximation, with category-level weight sharing enabling cross-OAR generalization from a minimal training set. Upon achieving full constraint compliance, the system optionally enters Lung Sparing Mode, in which the Lung RL agent continues to reduce lung dose through single-step LLM-supervised iterations. The two hybrid configurations utilizing the proposed framework were evaluated on 62 retrospective LA-NSCLC cases against an institutional KBP model and an LLM-only system using institutional dose-volume constraints. RESULTS: All three LLM-guided systems achieved a 98% clinical goal achievement rate, compared with 74% for KBP alone. Among the 16 cases requiring post-KBP refinement, the hybrid framework reached the clinical goal with 75% fewer LLM supervisory calls than the LLM-only system (1.7 ± 0.7 vs 6.7 ± 5.7) and in less time (20.6 ± 15.7 vs 33.4 ± 32.1 min), with the worst case reduced from 109.8 to 58.5 min. Enabling Lung Sparing Mode further reduced lung V20 (21.3 ± 8.4% vs 25.2 ± 10.1%, p < 0.001) and mean lung dose (14.2 ± 4.4 vs 15.0 ± 4.8 Gy, p < 0.001) relative to Normal Mode, at the cost of small increases in spinal cord and esophagus maximum dose that remained well within constraints. The resulting plans achieved lower lung dose than the clinically delivered plans for the same patients (V20 21.3% vs 24.6%; mean 14.2 vs 14.9 Gy). CONCLUSIONS: Integrating modular RL optimization with LLM-based supervision improves both lung sparing and planning efficiency relative to LLM-only automated planning, while enabling scalable deployment through cross-OAR weight sharing with reduced training requirements.

Abstract as published, via PubMed.

View on PubMedFull text at the publisherOpen in the app

For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.