Skip to main content
Back to timeline
Health Care Management ScienceSource publication:

Counterfactual Prescriptions via Hierarchical ML for Missed Chemotherapy Appointment Prevention

Synopsis

Using 1,825,948 chemotherapy appointment records from the Dana-Farber Cancer Institute, this study builds a hierarchical machine learning pipeline that first predicts cancellations and then no-shows, achieving F1-scores of 0.76 and 0.82 and improving minority-class performance by 7–10 points over a single-stage multinomial baseline; it further uses semi-supervised learning to infer no-show reasons from short-notice cancellations (weighted F1 of 0.57 and 0.54) and applies counterfactual simulation to evaluate interventions, finding that standard reminders are less effective than previously reported while provider consistency and commitment-based scheduling can reduce cancellations and no-shows.

Source-provided article image: Counterfactual prescriptions via hierarchical ML for missed chemotherapy appointment prevention
Figure 3

Figure 3.1: Distribution of the number of visits per patient

· Page 23

Interpretation

A hierarchical two-stage design (first separating cancellations, then separating shows from no-shows) outperforms a single-stage multiclass baseline on minority classes. Relative to prior work that conflates cancellations and no-shows or predicts only one outcome, this study decomposes the two missed-visit events into sequential binary problems and compares them directly against single-stage models under the same architectures. Across 1,825,948 chemotherapy appointments, cancellation F1 rose from 0.78 to 0.87 (Decision Tree), 0.85 to 0.92 (XGBoost), and 0.83 to 0.92 (MLP); no-show F1 rose from 0.23 to 0.30, 0.13 to 0.18, and 0.04 to 0.12; show-class F1 remained stable at 0.94–0.95.

Combining static and temporal features, and enriching raw features with derived variables, improves minority-class detection. Prior studies often emphasize either static or temporal variables; this work directly compares feature families and variable types through ablation experiments. Ablation shows static + temporal reaches no-show F1 of 0.69 versus 0.12 for static only and 0.03 for temporal only; raw + derived reaches no-show F1 of 0.74 versus 0.69 for raw only and 0.11 for derived only; SHAPIQ attributes 24.6% of predictive contribution to static–temporal interactions.

Semi-supervised pseudo-labeling transfers reason labels from short-notice cancellations to no-shows, enabling reason prediction for both missed-visit types. Existing models typically predict only whether an appointment will be missed, and no-show reasons are rarely documented in scheduling systems; this study propagates labels from 77,702 cancellations within 24 hours to 25,377 no-shows. At a pseudo-labeling threshold of 0.60, XGBoost achieves accuracy 0.71 and weighted F1 of 0.70 for no-show reasons; final weighted F1-scores are 0.57 for no-shows and 0.54 for cancellations.

Counterfactual simulation indicates that common reminder-based interventions offer limited or negative benefit in this chemotherapy setting, while provider consistency and commitment-based scheduling show more promise. Prior work often stops at predictive accuracy or retrospective evaluation; this study estimates the expected impact of multiple interventions before deployment using the predictive model. SMS reminders change cancellations by −0.16% and no-shows by −0.43%; reduced lead time worsens no-shows by −1.80%; provider consistency improves cancellations by +2.19% but worsens no-shows by −1.92%; commitment-based scheduling improves both cancellations (+0.13%) and no-shows (+0.11%), the only intervention to do so.

Perspective

The framework targets outpatient chemotherapy scheduling and relies only on operational fields available in the scheduling system, so it applies to information schedulers and care coordinators can access at booking or rescheduling; the authors state that core elements such as hierarchical modeling, reason prediction, and counterfactual analysis may transfer to other outpatient settings such as primary care, mental health, and chronic disease management, and recommend validating accuracy and robustness in experimental settings.

Counterfactual results come from offline simulation rather than prospective experiments, and the authors explicitly do not claim causal interpretation; the data lack demographic and socioeconomic variables and exclude unstructured sources such as clinical notes; no-show reason labels are inferred through semi-supervised pseudo-labeling, so their reliability depends on how well reason distributions transfer from short-notice cancellations to no-shows; no-show F1 varies between 0.74 and 0.82 across data splits, suggesting sensitivity to the splitting strategy.

Sources