Computing loss on scaled targets cuts MASE in all 24 architecture-benchmark comparisons of time series foundation model pretraining, by 18.8% on GIFT-Eval on average
Synopsis
The work proves that reversing an affine scaling (e.g., ReVIN) before computing the loss multiplies each series' gradient by a power of the scaling denominator, letting high-scale series dominate training (ScaleCon); computing a homogeneous residual loss directly on scaled targets (ScaleIn) instead leaves every mini-batch gradient and the full optimization trajectory invariant to arbitrary independent positive rescaling of training series, and in pretraining across four TSFM architectures it lowers MASE in all 24 architecture-benchmark comparisons, by 18.8% on GIFT-Eval and 21.9% on the M-competitions on average.
Figure 1: ScaleIn removes scale-driven gradient bias (Theorem 3.2 ). (a) Synthetic series scales span b ∈ { 1 , 4 , 13 , 52 } b\in\{1,4,13,52\} . (b) Under ScaleCon , MSE gradients with respect to z ^ \hat{z} scale as b 2 b^{2} , giving b = 52 b{=}52 approximately 2,700 × 2{,}700\times the gradient magnitude of b = 1 b{=}1 for equal scaled residuals. (c) ScaleIn leaves gradients unchanged by rescaling, on the same y y -axis. (d) Gradient distributions under ScaleCon (solid) and ScaleIn (hatched). (e) Growing scale induces gradient drift under ScaleCon but not ScaleIn . Lower-row curves use a 15-step moving average of | ∇ ℒ | |\nabla\mathcal{L}| .
arXivInterpretation
Reversing the scaling before computing the loss injects an implicit importance weight proportional to series scale into every gradient, causing high-scale series to dominate training. Prior empirical work observed that computing loss on scaled targets improves accuracy and hypothesized equal weighting of low- and high-scale series; this work formally characterizes the mechanism for degree-k homogeneous residual losses, deriving the scale-dependent gradient weights. Lemma 3.1 gives the gradient relation between ScaleCon and ScaleIn objectives, holding also for subgradients at nondifferentiable residuals; proof in Appendix B.
Computing the loss on scaled targets leaves every mini-batch gradient and the full optimization trajectory invariant to arbitrary independent positive rescaling of training series. Formalizes loss space, a previously rarely discussed design choice, as a scale-invariance condition, with applicability conditions for scale-equivariant scalers (standardization, min-max, robust scaling). Theorem 3.2 proves trajectory invariance under fixed initialization, mini-batch sequence, optimizer state, and optimizer randomness; Appendix B.1 extends to multivariate and quantile forecasting, Appendix B.2 discusses Chronos2's arcsinh nonlinearity.
Adopting ScaleIn in TSFM pretraining lowers MASE in all 24 architecture-benchmark comparisons and wins WQL in 21 of 24. Paired pretraining runs on Chronos2, Moirai-2.0, TimesFM2.5, and PatchTST with identical data order, initialization, and optimization settings, evaluated across M1, M3, M4, Tourism, Favorita, and GIFT-Eval. MASE reductions average 10.2% to 24.1% across architectures, WQL gains 3.3% to 19.1%; 18.8% average reduction on GIFT-Eval and 21.9% on the M-competitions; consistent across four different sources of scaling statistics.
The gains extend to supervised neural forecasting and masked reconstruction, with a minimal code change. Across NHITS, NBEATS, PatchTST, TCN, and LSTM, MASE falls in 16 of 20 matched comparisons and WQL in 15 of 20; MOMENT's median reconstruction MASE also drops. Supervised experiments pool M1, M3, M4, and Tourism series by frequency, with 4 random seeds; at the frequency level MASE falls in 49 of 75 comparisons and WQL in 43 of 75, concentrated in yearly and quarterly pools; MOMENT results averaged over 4 seeds.
Perspective
The result targets forecasting pipelines trained on heterogeneous datasets, especially cross-domain TSFM pretraining; it applies when the scaling method is scale equivariant and the loss is a homogeneous residual loss, covering standardization, min-max, and robust scaling, and MSE, MAE, and quantile loss. Multivariate forecasting with per-channel scaling, quantile forecasting, and masked reconstruction are all within scope. For practitioners, adoption means computing the loss on scaled targets using existing context statistics, a one-line change in most pipelines; predictions still return to the original scale for inference and evaluation. Fixed additive terms or scale floors make exact invariance approximate but do not remove the scale weighting in the gradient.
In multivariate settings with independent scaling of each series the scale bias persists, but its practical impact and the effects of more general normalization schemes that jointly transform correlated channels remain to be studied; how loss-space choices interact with probabilistic objectives and their parameterizations also lies beyond the homogeneous losses studied here. The supervised neural forecasting gains are broad but modest and vary between architectures and benchmarks, with frequency-level results concentrated in yearly and quarterly pools, and classical statistical references remain competitive on several benchmark-metric pairs. In addition, some numbers and table cells in the loaded text are missing after parsing, such as the surveyed model counts, some percentages, and several table entries, so restatements of those specific figures follow the readable text.
