Skip to main content
Back to timeline
arXivSource publication:

A Bayesian reconciliation framework cut cross-origin disagreement among overlapping multi-step forecasts by 58.5% and moved S&P 500 90% interval coverage from 94.8% to 88.3%

Related research and updates

Synopsis

The study proposes a model-agnostic Bayesian state-space post-processing framework that treats overlapping forecasts of the same future time issued from different origins as noisy, potentially biased measurements of a common latent trajectory, enforcing cross-origin coherence without access to the forecasting model's architecture, parameters, or training data; across 100 simulation replications it reduced cross-origin incoherence by a mean 58.5% (SD 15.2%) with mean latent-trajectory correlation 0.994, and in an S&P 500 volatility application it moved 90% prediction interval coverage from 94.8% to 88.3% and lowered CRPS by 8.7% while RMSE rose about 15%, with LSTM forecasts already showing low raw incoherence that reconciliation eliminated.

Source-provided article image: Bayesian Reconciliation of Overlapping Multi-Step Forecasts: A Post-Processing Framework for Cross-Origin Coherence

(a) Distribution of incoherence reduction. Mean: 58.5%, SD: 15.2%.

arXiv

Interpretation

The paper formalizes cross-origin incoherence as disagreement among forecasts for the same target time issued at different forecast origins, and defines a scalar metric as the average squared pairwise disagreement, which is zero for a perfectly coherent system. Existing forecast reconciliation work addresses coherence under hierarchical, temporal, and cross-temporal aggregation constraints, while stability work studies consistency as the origin changes; this work instead reconciles overlapping forecasts of the same target timestamp, a different constraint structure. The metric is defined in the text and illustrated with a forecast matrix (rows as origins, columns as target times); the authors state it is used as a post-processing diagnostic and do not claim it as a novel stability measure.

The paper develops a Bayesian state-space reconciliation framework in which the latent coherent trajectory follows a stationary AR(1) process and each forecast is a noisy observation with horizon-specific additive bias and horizon-specific variance, with posterior inference via NUTS/HMC in Stan. Relative to the closest methodological antecedent (the hierarchical Bayesian state-space reconciliation of Eckert et al. 2021), the latent state here is the coherent trajectory itself rather than a reconciliation error, the constraint is an identity constraint at each target time rather than a linear aggregation constraint, and horizon-specific bias and reliability parameters have no counterpart in hierarchical reconciliation. Model equations, likelihood factorization, and priors are given in the text; the authors resolve location non-identifiability with a sum-to-zero constraint and provide a derivation that the parameters are uniquely identified.

Across 100 independent simulation replications, reconciliation reduced incoherence by a mean 58.5% (SD 15.2%, range 24.7%–97.8%) with mean correlation 0.994 to the true latent state, and 95% of replications achieved correlations above 0.988. The simulation deliberately introduces nonlinear horizon-dependent noise, correlated errors across adjacent horizons, and occasional large deviations with 5% probability that violate model assumptions, yet reports stable performance, indicating usefulness beyond idealized conditions. 100 replications with 2 chains of 1,000 iterations (500 warm-up), convergence (R-hat) for all parameters in 99% of replications with no divergent transitions, effective sample sizes above 400 for key parameters, and an average 2.8 seconds per replication.

In the S&P 500 volatility application, reconciliation drove incoherence below numerical precision, moved 90% prediction interval coverage from 94.8% to 88.3% closer to nominal, lowered CRPS by 8.7%, and raised RMSE by about 15%; LSTM forecasts already had low raw incoherence and their remaining incoherence was eliminated. The results explicitly quantify the coherence-versus-point-accuracy trade-off and show the same post-processing layer applies to both a statistical benchmark and a neural forecasting system without modifying the underlying model. Uses SPY ETF data from December 2023 through December 2025, 5-day rolling volatility over 516 trading days, with the final 100 days held out for rolling out-of-sample evaluation; four chains of 1,000 iterations (300 warm-up), all parameters meeting R-hat targets with effective sample sizes above 500, and alternative priors changing reliability posterior means by less than 1%.

Perspective

The framework targets already-generated direct multi-step forecast matrices and suits settings where retraining the underlying model is impossible or costly, such as black-box or computationally expensive forecasting systems; it offers risk managers, supply chain planners, and epidemic forecasters a way to enforce cross-origin coherence after the fact and to obtain horizon-specific bias and reliability diagnostics without seeing model internals. The method is derived under a stationary AR(1) latent process and Gaussian observation assumptions, and both simulation and empirical evaluation stay within that setting; the authors position it as complementary to stability-aware training, which reduces origin disagreement during training while this framework enforces coherence after forecasts are generated.

In the nonlinear stress test, incoherence was fully eliminated but correlation with the true latent trajectory fell from 0.994 in the linear case to 0.633, showing that enforcing coherence and accurately recovering the latent process are related but distinct objectives that readers should keep separate. In the empirical study RMSE rose about 15% while CRPS fell 8.7%, and how this trade-off varies across series and forecast quality remains an open question. Horizon-specific bias and reliability are assumed time-invariant, which the authors note may be restrictive under regime shifts or model drift; the Gaussian observation assumption, conditional independence across forecasts, and support only for direct rather than recursive multi-step forecasts further delimit the current formulation. The authors also state that no separate numerical prior-sensitivity experiment was performed, leaving systematic testing of alternative prior families for future work.

Sources