Skip to main content
Back to timeline
arXivSource publication:

OncoAgent turns esophageal radiotherapy guideline text into 3D target volumes zero-shot, reaching CTV Dice 0.842 with no significant difference from fully supervised nnU-Net

Synopsis

The study introduces OncoAgent, a guideline-aware AI agent framework in which a large language model parses free-text clinical guidelines into an executable tool-call plan, generating three-dimensional target volumes without any expert-annotated training data; on planning CT from 40 mid-thoracic esophageal cancer patients (32 for training, 8 for testing), it achieved zero-shot CTV Dice of 0.842 and PTV Dice of 0.880, with no statistically significant difference from the fully supervised nnU-Net (GTV Prior) at 0.862 and 0.893 on primary metrics, and it was rated higher than that supervised baseline by blinded physicians on guideline compliance, modification effort, and clinical acceptability.

Source-provided article image: A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-delineation

Interpretation

A two-phase (planning and execution) agentic architecture converts free-text clinical guidelines into 3D target volumes in a completely training-free manner. Prior deep learning segmentation relies on expert-annotated data and requires re-annotation and retraining whenever guidelines change; this work instead has an LLM reason from guideline text and invoke pre-trained OAR segmentation models plus geometric operation tools (segment, dilate, subtract, and so on) to build contours. The method specifies three system-prompt strategies (clinician role assignment, guideline parameterization, robust tool-call sequencing) and fixed templates CTV = (GTV ⊕ m_ctv) \ OARs and PTV = CTV ⊕ m_ptv; plans are validated against a JSON schema and iteratively self-refined on violations. End-to-end execution averaged 1.74 minutes per patient with a mean of 1.13 LLM inference calls per case.

On esophageal cancer, zero-shot performance is comparable to a fully supervised state-of-the-art baseline. Earlier supervised esophageal target segmentation models (DDAU-Net, nnU-Net) struggled with invisible, guideline-derived CTV boundaries; adding GTV prior knowledge produced large gains, and OncoAgent reaches that same level zero-shot. 40 mid-thoracic esophageal cancer planning CT cases, 32 for training and 8 for testing; CTV Dice 0.842±0.026 with MSD 1.06±0.15 mm, PTV Dice 0.880±0.015 with MSD 1.12±0.14 mm, versus nnU-Net (GTV Prior) at CTV 0.862±0.026 and PTV 0.893±0.021, with paired t-tests showing p ≥ 0.05 on primary metrics such as DSC and MSD. OncoAgent had higher sensitivity (CTV 0.884 vs. 0.845; PTV 0.922 vs. 0.873) and slightly lower precision.

In blinded clinical assessment, physicians preferred OncoAgent over the supervised baseline. Evaluation moves beyond volumetric overlap to clinically meaningful dimensions: guideline compliance, modification effort, and clinical acceptability. Two senior radiation oncologists (17 and 20 years of experience) blindly rated the 8 test cases on a 5-point Likert scale: guideline compliance 4.06±0.68 vs. 3.56±0.51, modification effort 4.12±0.89 vs. 3.44±0.63, clinical acceptability 3.81±0.75 vs. 3.38±0.72; the share of ratings ≥4 was 81.2% vs. 56.2%, 68.8% vs. 37.5%, and 75.0% vs. 50.0%, respectively.

The framework transfers zero-shot to alternative esophageal guidelines and to other anatomical sites. Supervised models need re-annotation and retraining with each guideline update, whereas this framework adapts to different protocols simply by updating the textual input. Tool Call F1 was 1.00 for the IJROBP expert consensus guideline, 0.78 for JASTRO 2024, 0.73 for the CROSS trial, and 0.64 for prostate (RTOG 0126), the lower prostate score attributed to unfamiliar anatomical terms such as “SV_prox1cm”; all without task-specific fine-tuning.

Perspective

The result applies to CTV/PTV auto-delineation on planning CT for mid-thoracic esophageal cancer radiotherapy, in settings that have a pre-trained OAR segmentation model, can expose structured tool interfaces, and retain physician review. Its value lies in adapting to guideline updates or institutional protocol variation by replacing the textual input rather than re-annotating data or retraining models, and it can extend to other guidelines and anatomical sites (prostate is the example given). For a reader, this means the conversion from guideline text to an executable delineation plan can be treated as a separate, reviewable step rather than delegating full trust to an end-to-end black box.

Several points remain worth watching: first, performance depends on the quality of the underlying OAR segmentation model; second, as an LLM-based system it is susceptible to semantic misinterpretation and hallucination, for instance misapplying a guideline clause to an inappropriate anatomical context, and it cannot capture implicit clinical knowledge absent from the guideline text, so clinician review of the execution plan remains essential; third, strictly guideline-compliant contours may lack the pragmatic safety margins physicians manually draw to accommodate esophageal motion from respiration, swallowing, and cardiac pulsation, and future work should integrate adaptive margin strategies; fourth, the study is limited to 40 patients from a single institution, and multi-center validation, extension to other sites such as head and neck and cervical cancer, and clinician feedback loops remain open directions. In addition, this reading covered the full text with figures presented as text and tables, so readers wanting to inspect specific contour shapes should consult Figure 3 in the original.

Sources