Skip to main content
Back to timeline
arXivSource publication:

RadOnc-Agent splits radiotherapy into four phases and 26 callable functions, hitting the intended function in 98.79% of single-function executions

Synopsis

The work presents RadOnc-Agent, an agentic framework that formalizes radiotherapy into four clinical phases and exposes 26 callable functions through a conversational interface, with a large-language-model controller mapping clinical intent to schema-constrained calls while preserving patient and workflow context; across 2,600 single-function requests, 200 prespecified synthetic cross-stage scenarios, and 120 workflow instances from 60 de-identified patient records, it selected the intended function in 98.79% of single-function executions, completed 96.50% of scripted cross-stage workflows, and completed 96.67% of real-patient workflow executions.

Source-provided article image: RadOnc-Agent: An LLM-Orchestrated Framework for AI Workflows Across the Radiotherapy Care Pathway

[Uncaptioned image]

arXiv

Interpretation

The framework formalizes the longitudinal radiotherapy workflow into four clinical phases and exposes 26 callable functions through a conversational interface, so capabilities scattered across clinical stages, software environments, and data modalities can be coordinated. Prior AI work advanced individual radiotherapy tasks that remained separated; this work makes the longitudinal pathway from treatment decision-making through follow-up the object of orchestration rather than a single task. Descriptive evidence from the four-phase formalization and the 26-function interface design, paired with the execution evaluations.

A large-language-model controller maps clinical intent to schema-constrained calls, preserves patient and workflow context, and routes requests to specialist services. Schema constraints, identity validation, and longitudinal state are treated as explicit architectural components rather than relying on unconstrained model-generated calls. Ablations show that removing longitudinal state reduced cross-stage completion from 96.50% to 84.00%, and disabling schema and identity validation increased mismatched backend dispatch from 0% to 95.28% in a replay/test evaluation.

System execution was evaluated across single-function, synthetic cross-stage, and real-patient workflow settings, with high reported completion rates. The evaluation spans 2,600 single-function requests (7,800 repeat executions), 200 prespecified synthetic cross-stage scenarios (600 executions), and 120 workflow instances from 60 de-identified patient records (360 clean executions), covering decision-to-planning and planning-to-adaptation. Intended-function selection was 98.79% for single-function executions, 96.50% completion for scripted cross-stage workflows, and 96.67% completion for real-patient workflow executions.

The authors explicitly bound the conclusion: these findings establish the technical feasibility of an LLM-orchestrated architecture for coordinating heterogeneous radiotherapy capabilities and information. Feasibility is separated from clinical correctness, clinical utility, and prospective benefit, so execution success is not equated with clinical effect. The text states directly that the findings do not establish clinical correctness, clinical utility, or prospective benefit.

Perspective

The results target coordination of heterogeneous capabilities and information in radiotherapy, applying to longitudinal workflows from treatment decision-making through follow-up, including decision-to-planning and planning-to-adaptation. For a reader, the value lies in a reusable architectural pattern: formalize a domain workflow into phases, wrap capabilities as callable functions, let an LLM controller perform schema-constrained intent mapping and routing, and maintain context through longitudinal state. This pattern can be borrowed for other clinical or engineering workflows that need cross-stage coordination, for prototyping and execution-level validation.

The available text is summary-level: it does not give implementation details of each function, the specific rules for schema and identity validation, the representation of longitudinal state, the basis for constructing synthetic scenarios, or the selection criteria for real-patient workflow instances. The relationship between execution success rates and clinical correctness remains to be studied, and differences between the replay/test evaluation setting in the ablation and real deployment conditions deserve attention. These are scope and open questions rather than shortcomings of the work.

Sources