Skip to main content
Back to timeline
Journal of pediatric hematology/oncologySource publication:

AI-Assisted Generation of Long-Term Follow-Up Recommendations for Survivors of Childhood Cancer and Hematopoietic Stem Cell Transplantation

Synopsis

This feasibility study fed deidentified treatment summaries into OpenAI GPT-4o with structured prompts to draft long-term follow-up recommendations for childhood cancer and HSCT survivors, then compared them with clinician-generated recommendations based on institutional standards and Children's Oncology Group Long-Term Follow-Up Guidelines (version 5) in an independent validation cohort of 40 survivors, finding 467 AI-generated versus 446 clinician-generated items, with 385 of 446 clinician recommendations (86.3%) also identified by AI and 385 of 467 AI recommendations (82.4%) also present in clinician plans, while most discordant recommendations involved radiation-related exposures and survivorship scenarios requiring nuanced clinical interpretation.

Source-provided article image: AI-Assisted Generation of Long-Term Follow-Up Recommendations for Survivors of Childhood Cancer and Hematopoietic Stem Cell Transplantation.
FIGURE 1 ·

FIGURE 1 Open multimedia modal Study workflow for AI-generated survivorship LTFU recommendations.

PubMed

Interpretation

The study shows a large language model can generate draft long-term follow-up recommendations with substantial concordance with clinician-generated recommendations under structured prompting. Prior studies explored AI-assisted clinical recommendations, documentation, and care planning in other health care settings, but data on LLMs for pediatric oncology survivorship care planning were limited; this study applies GPT-4o in that setting and reports bidirectional agreement. In an independent validation cohort of 40 survivors, 385 of 446 clinician recommendations (86.3%) were also identified by AI, and 385 of 467 AI recommendations (82.4%) were also present in clinician plans; the cohort included chemotherapy 100%, radiation 27.5%, and HSCT 40%.

Discordant recommendations mainly involved radiation-related exposures and survivorship scenarios requiring nuanced clinical interpretation. The study lists specific discordance categories: among clinician-only recommendations, the most frequent were neuropsychological testing (10), lipid screening (8), electrocardiography (6), echocardiography (5), and clinical breast examinations (5); among AI-only recommendations, the most common were urinalysis (17), renal and bladder ultrasound (14), colonoscopy (7), vitamin D monitoring (6), and bone density evaluation (5). Discordant recommendations were reviewed by a survivorship nurse with expertise in childhood cancer long-term follow-up care to confirm categorization, but adjudication was not blinded, no second reviewer independently assessed discordant recommendations, and inter-rater agreement was not calculated.

The study positions AI-generated recommendations as an initial draft for clinician-supervised workflows rather than a final survivorship care plan. The authors note the principal risk is omission of an appropriate recommendation rather than inclusion of an inappropriate one, so they advise clinicians to verify treatment exposures, carefully review radiation-related recommendations, confirm institution-specific practices, and incorporate patient-specific considerations. This positioning comes from the discussion and conclusion of this feasibility study and is a workflow suggestion based on concordance results, not direct evidence from clinical outcomes or efficiency metrics.

Perspective

This work applies to generating draft long-term follow-up recommendations for childhood cancer and HSCT survivors within a structured, guideline-informed framework for clinician-supervised use; its prompts were based on institutional survivorship practices and COG Long-Term Follow-Up Guidelines version 5, the validation cohort was 40 survivors at a single institution, and AI recommendations were evaluated after clinician-prepared treatment summaries had already been created, so it does not address automation of treatment abstraction from electronic medical records.

Readers should still watch that discordant recommendations were adjudicated by a single unblinded reviewer who prepared most of the clinician-generated plans, and inter-rater agreement was not assessed; patients were selected purposively rather than by random or consecutive sampling; the validation cohort included few CNS tumor survivors, and because CNS tumor survivors frequently receive cranial radiation, the exposure category with the greatest AI discordance, the observed agreement may overestimate performance in a cohort with a higher proportion of CNS tumor survivors; and revalidation will be necessary whenever the underlying model or prompting strategy is modified.

Sources