A unified scaling law predicts held-out Toto 2.0 capacity curves at 1.09% and 1.50% MAPE, and activation interventions show frozen time series foundation models reuse historical rules across queries
Related research and updatesSynopsis
Using 18,768 experimental cells from 21 checkpoints on 23 dataset-frequency tasks across six domains, the work fits a five-parameter unified scaling law linking capacity, input length, and forecast horizon, and predicts the fully held-out Toto 2.0 series (about 4M-2.5B parameters) with MAPEs of 1.09% at input length 2048 and 1.50% at 4096 on horizon-averaged capacity curves; it also proposes a unified theory of time series learning in which Gaussian regression, matched-history comparisons, parameter exchanges, and activation interventions support the account that full-shot models accumulate rule information in weights while frozen time series foundation models extract and reuse historical rules through activations.
Figure 1: The Unified Scaling Law and observed resource profiles. (a) Predicted MASE over N , L N,L at H = 48,192,720 H=48,192,720 (blue, green, orange). (b)–(d) Observed means (markers) and mean predictions (lines) over identical capacity, input-length, and horizon groups, using all 1,070 Controlled points. Resource axes are logarithmic; R R is linear in (a) and logarithmic in (b)–(d). 1k denotes 1,024 steps.
arXivInterpretation
A five-parameter unified scaling law places pretrained capacity, input length, and the scored forecast horizon on one response surface: capacity gains increase with history, context gains diminish toward saturation, and horizon effects enter as a common shift. Prior work established training-resource scaling, capacity-lookback interactions, or long-context benefits separately, but left their joint description across heterogeneous released checkpoints unresolved; this work constructs and estimates the law from local resource relations under a parsimony prior. Aggregated from 18,768 experimental cells across 21 checkpoints, 23 dataset-frequency tasks, and six domains into 1,070 controlled fitting points; MAPE is 10.13% on the 1,070 controlled points and 12.82% on 111 non-overlapping GIFT Public points.
The law remains predictive on a fully held-out architecture: after removing all five Toto 2.0 checkpoints and re-estimating, prediction on its 150 controlled grid points gives 11.53% MAPE, and horizon-averaged capacity curves give 1.09% MAPE at input length 2048 and 1.50% at 4096. Earlier scaling studies largely report fit within the model families used for estimation; this work tests transferability across heterogeneous designs via a held-out architecture. The leave-out procedure removes every Toto 2.0 observation before reconstructing local targets and re-estimating all five coefficients; among reference forms, constant prediction gives 17.10% and the capacity-independent form 11.84%, versus 11.53% for the unified law.
A unified theory of time series learning proposes that full-shot models learn by accumulating information in weights while frozen time series foundation models use history by extracting information through activations; matched-history comparisons establish the predictive value of additional history, and parameter exchanges show rule information can be retained and applied to new queries. The in-weight versus in-context distinction is carried into forecasting and separated into controlled tests of rule identification, retention, and reuse. On 23 tasks, increasing history from 1,024 to 8,192 reduces DLinear's MASE from 1.042 to 0.931 and PatchTST's from 0.885 to 0.833; parameter exchange moves predictions toward the donor rule in all 54 process-architecture-length-seed conditions.
Activation interventions indicate that history-derived rule information in a frozen model can be transferred across queries and can recover a contribution to natural long-context prediction: transfer reduces mean error by 7.51%, and recovery restores 41.67% of the mean error increase caused by erasure. Prior analysis of time series foundation models largely compared errors; this work uses activation patching to test directly how historical states affect independent queries. In TimesFM 2.5, erasure raises error while transfer and recovery improve it with positive paired-query 95% intervals in all 18 conditions; matching-rule donors outperform opposite-rule donors in every tested condition across three layers and three strengths, and Chronos 2 recovery lowers mean error by 6.01% on independent confirmation data.
Perspective
The law describes average resource returns within the observed ranges of capacity, input length, and horizon: the controlled input grid ends at 8192, long-context support comes from nine of 21 checkpoint interfaces, and 12 of the 23 dataset-frequency tasks provide at least 8192 preceding observations at every scored window. The fitted context saturation scale is roughly 3.6k-4.6k over 10M-1B parameters, and mean paired gains become slightly negative beyond 4096, so the law most directly serves practitioners choosing model size, input length, and horizon under deployment latency and memory constraints, and model developers planning pretraining scale and validation coverage. The learning-theory portion concerns three stationary synthetic processes and the frozen TimesFM 2.5 and Chronos 2, applying to settings where the rule is identifiable from a historical prefix and query and history marginal distributions are fixed.
Several open questions remain: how the law behaves beyond input length 8192, outside the observed capacity range, and on domains not in the panel; whether the capacity dependence of the saturation scale holds for larger models; whether the layer, token span, and strength selected for TimesFM 2.5 transfer to other architectures, given that Chronos 2 direct transfer leaves overall error nearly unchanged while recovery still helps; and how far the synthetic-process conclusions extend to real non-stationary series. In addition, several key expressions appear as equation images and their symbol-level details are not restated in the prose, so the precise functional forms of the local elasticities and the saturation scale can only be understood from the surrounding description.
