Public articles linked to the same research event.
arXiv Using 18,768 experimental cells from 21 checkpoints on 23 dataset-frequency tasks across six domains, the work fits a five-parameter unified scaling law linking capacity, input length, and forecast horizon, and predicts the fully held-out Toto 2.0 series (about 4M-2.5B parameters) with MAPEs of 1.09% at input length 2048 and 1.50% at 4096 on horizon-averaged capacity curves; it also proposes a unified theory of time series learning in which Gaussian regression, matched-history comparisons, parameter exchanges, and activation interventions support the account that full-shot models accumulate rule information in weights while frozen time series foundation models extract and reuse historical rules through activations.
Using 18,768 experimental cells from 21 checkpoints on 23 dataset-frequency tasks across six domains, the work fits a five-parameter unified scaling law linking capacity, input length, and forecast horizon, and predicts the fully held-out Toto 2.0 series (about 4M-2.5B parameters) with MAPEs of 1.09% at input length 2048 and 1.50% at 4096 on horizon-averaged capacity curves; it also proposes a unified theory of time series learning in which Gaussian regression, matched-history comparisons, parameter exchanges, and activation interventions support the account that full-shot models accumulate rule information in weights while frozen time series foundation models extract and reuse historical rules through activations.
Using 18,768 experimental cells from 21 checkpoints on 23 dataset-frequency tasks across six domains, the work fits a five-parameter unified scaling law linking capacity, input length, and forecast horizon, and predicts the fully held-out Toto 2.0 series (about 4M-2.5B parameters) with MAPEs of 1.09% at input length 2048 and 1.50% at 4096 on horizon-averaged capacity curves; it also proposes a unified theory of time series learning in which Gaussian regression, matched-history comparisons, parameter exchanges, and activation interventions support the account that full-shot models accumulate rule information in weights while frozen time series foundation models extract and reuse historical rules through activations.
Using 18,768 experimental cells from 21 checkpoints on 23 dataset-frequency tasks across six domains, the work fits a five-parameter unified scaling law linking capacity, input length, and forecast horizon, and predicts the fully held-out Toto 2.0 series (about 4M-2.5B parameters) with MAPEs of 1.09% at input length 2048 and 1.50% at 4096 on horizon-averaged capacity curves; it also proposes a unified theory of time series learning in which Gaussian regression, matched-history comparisons, parameter exchanges, and activation interventions support the account that full-shot models accumulate rule information in weights while frozen time series foundation models extract and reuse historical rules through activations.