In simulations built on two randomized controlled trial datasets, Cox proportional hazards and random survival forest performance diverged by performance measure and proportional hazards assumption
Synopsis
The authors conducted a comprehensive neutral simulation comparison based on two reference datasets from randomized controlled trials, evaluating the Cox proportional hazards model and random survival forest for patient-specific survival probability prediction across multiple performance measures following TRIPOD recommendations, and found that conclusions based solely on the C index may not generalize to other aspects of predictive performance, that measures of overall performance may generally give more reasonable results, that the standard log-rank splitting rule for the random survival forest may be outperformed by alternative splitting rules particularly in nonproportional hazards settings, that the random survival forest performance suffered less in data with treatment-covariate inte
Interpretation
The study built multiple simulation scenarios based on two reference datasets from randomized controlled trials to neutrally compare the Cox proportional hazards model and random survival forest, using multiple performance measures according to TRIPOD recommendations. Previous studies comparing these two methods were predominantly based on real-world observational time-to-event data and relied mainly on the C index as a single measure; this work places the comparison in a simulation framework grounded in randomized controlled trial settings and extends it to multiple performance dimensions. Evidence comes from a simulation study based on two reference datasets from randomized controlled trials, a design suited to methodological comparison; the abstract does not report the number of simulation scenarios, sample sizes, or specific effect sizes.
Conclusions based solely on the C index may not be generalizable to other aspects of predictive performance, and measures of overall performance may generally give more reasonable results. This directly addresses methodological criticism of the C index as the dominant comparison measure, showing that a single measure's ranking does not represent overall model performance. The conclusion comes from comparisons among different performance measures within the simulations; the abstract does not report specific values or magnitudes of differences across measures.
The standard log-rank splitting rule used for the random survival forest may be outperformed by alternative splitting rules, particularly in nonproportional hazards settings. This indicates that implementation details of the random survival forest, specifically the choice of splitting rule, can themselves affect its relative performance, beyond algorithm-level differences. Evidence comes from comparisons of different splitting rules within the simulations; the abstract does not specify the alternative splitting rules or the magnitude of improvement.
In the simulations, the random survival forest performance suffered less in data with treatment-covariate interactions, while the Cox proportional hazards model performance was affected by violation of the proportional hazards assumption. This identifies data settings in which each method is more suitable or more vulnerable, providing conditional guidance for method choice in randomized controlled trial analyses. The conclusion is based on simulation settings involving interaction terms and the proportional hazards assumption; the abstract does not provide quantified performance differences.
Perspective
This work is aimed at researchers and statistical method users who predict patient-specific survival probabilities from randomized controlled trial time-to-event data, helping them make conditional choices between the Cox proportional hazards model and random survival forest and highlighting reporting of multiple performance measures per TRIPOD recommendations. Its conclusions apply to the scenarios set up in these simulations, particularly those involving nonproportional hazards and treatment-covariate interactions; applicability to other data sources, follow-up structures, or disease contexts requires separate evaluation.
Because the loaded text is incomplete, containing only the abstract, keywords, and references, without the main text's simulation scenario settings, performance measure values, splitting rule details, or figures and tables, specific effect sizes, scenario counts, or rankings under each measure cannot be summarized here. Readers who wish to adjust their own analysis workflows accordingly should consult the original methods and results sections to confirm how well the simulation conditions match their data, and should pay attention to the specific implementation of alternative splitting rules.
