Researchers at Mayo Clinic Arizona have built an agentic AI framework intended to connect specialist tools across the full radiation-therapy pathway, from treatment decisions and planning through adaptation and follow-up. The system, RadOnc-Agent, uses a large-language-model controller to translate clinical requests into structured calls to 26 functions. In a new arXiv preprint, the team reports high rates of correct tool selection and technical workflow completion, while emphasizing that the study does not demonstrate clinical correctness, clinical utility or improved patient outcomes.
Radiotherapy is a chain of interdependent decisions rather than one isolated procedure. A prescription affects which anatomical structures must be outlined; imaging and beam geometry constrain dose planning; changes observed during treatment can trigger reassessment; and delivered dose informs later toxicity and recurrence monitoring. Existing AI systems often address one step at a time, leaving people to identify the right model, assemble its inputs and carry results between separate software environments.

RadOnc-Agent is designed as a coordination layer rather than a single model that performs every medical task. Its 17 core functions cover four phases: diagnosis and treatment decisions, simulation and planning, on-treatment adaptation, and follow-up and outcomes. Nine additional utilities handle tasks such as patient-status retrieval, imaging and implant lookup, clinical-trial searches, schedules, messages, literature searches and deterministic calculations. Specialist services perform imaging, dose and quantitative work, while the language model manages intent, routing and workflow state.
The researchers evaluated 2,600 unique single-function requests, repeating each request three times in fresh sessions for 7,800 executions. The controller selected the intended function in 98.79% of those runs, produced schema-conforming calls in 99.27% and returned parseable structured output in 98.64%. Deliberately confusable requests were harder than direct ones: tool-selection accuracy was 96.73% for the confusable set versus 99.79% for explicit requests.
A second benchmark tested 200 prespecified synthetic scenarios spanning multiple clinical stages, again repeated three times. RadOnc-Agent completed all required calls in 96.50% of the 600 executions and followed the reference trajectory in 95.00%. Removing persistent workflow state reduced completion to 84.00%, suggesting that carrying patient and task context between steps was central to the result. The full-pathway scenarios were the most demanding, with completion falling to 92.67% as call chains grew longer.

The team also ran a retrospective, offline evaluation using 60 de-identified patient records that had been held out from framework development and specialist-model training or tuning. Each record contributed a decision-to-planning workflow and a planning-to-adaptation workflow, producing 120 workflow instances and 360 clean repeat executions. The system completed all required functions in 96.67% of those executions and preserved context throughout 98.61%. No output was used to direct patient care.
The paper highlights deterministic checks as an essential safety layer. Before dispatch, the system validates function schemas and patient, course and plan identifiers. With those gates enabled, no mismatched call was dispatched in the identifier-integrity test. When the gates were bypassed, 343 of 360 mismatched calls were dispatched in the replay or test setting. The language model’s own behavior was not sufficient by itself: in a missing-input benchmark, it inappropriately tried to proceed in 44 of 1,560 executions.
Those results leave a large gap between technical orchestration and clinical adoption. The real-patient evaluation was retrospective, offline and conducted at one site; it did not test treatment decisions, delivery, clinician workload, automation bias or patient outcomes. The 26 functions also have different evidence bases and cannot be combined into one measure of clinical performance. The authors call for external retrospective validation, prospective silent deployment and then human-in-the-loop studies with mandatory professional review before broader use. RadOnc-Agent therefore represents a systems blueprint for connecting medical AI tools, not an autonomous radiation-oncology system ready for unsupervised care.

Comments
Loading comments…