Skip to main content
Back to timeline
发表出处待核验Source publication:

XGBoost beat NeuralProphet for Budapest metro M4 demand forecasting, and its SHAP explanations matched expert expectations

Synopsis

Using one year of hourly boarding data from Budapest metro line M4 (216 usable days, 05:00–24:00) restricted to a single station and direction (Kálvin tér toward Kelenföld vasútállomás) and a 6-hour forecast window, the study fairly compared inherently interpretable NeuralProphet with a black-box XGBoost interpreted post hoc via SHAP, finding XGBoost more accurate (MAE 91.0753 vs 230.2803, RMSE 144.7094 vs 363.2115, R-squared 0.9600 vs 0.7896) and SHAP temporal-feature insights consistent with NeuralProphet seasonality components and transport expert knowledge.

Source-provided article image: Practical perspectives on machine learning interpretability for public transport demand forecasting
Fig. 1

Fig. 1. Route and stations map of the metro line M4 in the city of Budapest.

· Page 2

Interpretation

On a real public transport demand forecasting task, black-box XGBoost clearly outperformed inherently interpretable NeuralProphet in accuracy. Both model families had separately shown success in transport time series, but their relative usefulness for transparent demand forecasting was less well understood; this work compares them under one dataset, one forecast window, and one evaluation pipeline. Reported MAE, RMSE, and R-squared on an unseen test subset of M4 hourly passenger data, averaged over the 6-hour prediction window, plus a qualitative comparison on a one-week continuous test window.

SHAP explanations of XGBoost reproduced the hourly pattern of a weekday morning peak around 08:00, a stronger afternoon peak around 16:00–17:00, and a flatter weekend distribution, aligning with dataset-wide average passenger counts and NeuralProphet seasonality components. Moves the usefulness of post-hoc interpretation from an abstract claim to concrete temporal patterns that can be checked against expert knowledge, and identifies the hour feature as dominant for XGBoost by magnitude. Compares SHAP values with NeuralProphet seasonality components using dataset-wide average passenger numbers as the reference for expert knowledge; the authors note the three rows represent different quantities yet often align.

XGBoost shifts its reliance from calendar features to recent observed lagged passenger counts as the day progresses, whereas NeuralProphet does not show the same lag-based behavior. Provides an interpretable clue about the black-box model's internal mechanism: with zero lagged values at night the model leans more on the hour itself, and after 10:00 lagged inputs already explain part of the variation, reducing the current hour's marginal contribution. Based on the pattern of SHAP values across hours, interpreted together with the fact that lagged passenger values are zero during night hours in this dataset.

The study argues for prioritizing predictive performance in public transport demand forecasting and holds that accurate models can still be viably interpreted with SHAP. Turns the accuracy-versus-interpretability trade-off into an actionable selection recommendation, while noting model-based approaches' advantages in nuanced interpretation, expert fine-tuning, and robustness, and proposing a hybrid direction. Conclusions rest on an empirical comparison for a single line and a single station-direction; the authors frame the recommendation as based on this study and note NeuralProphet deviates only slightly from classical Prophet, with model-based techniques such as deep unfolding or variable projections potentially closing the gap.

Perspective

This work speaks to transport planners and operators who must trade off accuracy against transparency, and it applies to short-term hourly demand forecasting settings with strong daily and seasonal cycles, especially commuter-oriented lines like M4 that do not operate at night. Its actionable message is that under similar data conditions one may favor the more accurate XGBoost and use SHAP to explain its decisions to stakeholders, while model-based approaches retain value for nuanced interpretation, expert fine-tuning, and robustness, making hybrid routes worth pursuing. The authors also extend the implications to micromobility, traffic volume forecasting, demand-responsive transit, and to electricity load forecasting and environmental monitoring with strong seasonal effects.

Passenger numbers come from algorithmic estimation via underground cellular antennas, a process prone to error that produces missing data, leaving only 216 days from 05:00 to 24:00, which may affect how stable the conclusions are under other data-quality conditions. The analysis is limited to a single station and direction, a simplification the authors made to control spatial heterogeneity, so generalization across stations and directions still needs testing. Evaluation uses dataset-wide average passenger counts as the reference for expert knowledge rather than an independent expert annotation process. In addition, the available text is an incomplete version, so the specific values and graphical details of Figures 1 to 3 cannot be verified, and the quantitative extent of the interpretability comparison can only be judged from the prose.

Sources