NanoForecast v0.5 fixes only its training pipeline, cutting MASE from 3.030 to 1.704 and beating 200M-parameter TimesFM on all three ETT datasets
Synopsis
Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
Interpretation
The paper documents three training pipeline defects: the data pipeline unconditionally included a horizon key so that with multi_horizon=False the point-forecast loss still backpropagated through the full context-length output; the multi-task loss computed the quantile term before truncating predictions to the forecast horizon, causing shape mismatches and incorrect gradient flow in the quantile branch; and v0.3 augmented windows only with scale, shift, and jitter, leaving training diversity thin. Prior work tends to attribute accuracy gaps to architecture or scale; here the three silent pipeline errors are named individually with fixes: truncate all loss terms to the forecast horizon, and augment each window in-loop with stochastically sampled jitter, random scaling, shifting, masking, and time reversal. The authors train v0.3 and v0.5 with the identical architecture, corpus, multi-task loss family, and compute budget, and compare them under the fixed protocol of Section 6.1, making this a controlled paired comparison; however, one checkpoint per configuration is released and seed-to-seed variance is not reported.
Fixing the pipeline alone, with no architecture change, cuts overall MASE from 3.030 to 1.704, a 43.8% reduction, improving all six datasets; exchange rate moves most (11.758 to 4.317), then ETTh2 (1.328 to 1.110), while ETTm1 (0.288 to 0.287) and ETTh1 (0.681 to 0.676) are ties for practical purposes. The gain comes from the training process rather than model structure or parameter count, which leads the authors to raise a field-level question: how many published gaps come from training setups rather than architectures. The ablation table gives per-dataset MASE for v0.3 and v0.5 across six datasets, and the authors state the released code reproduces the comparison; v0.3 is the validation-best snapshot at epoch 147 and v0.5 at epoch 51, with validation loss 0.2230 versus 0.2204.
At 6.5M parameters, v0.5 beats 200M-parameter TimesFM on four of six benchmarks: ETTh1 0.676 vs. 0.705, ETTh2 1.110 vs. 1.360, ETTm1 0.287 vs. 0.545, and exchange rate 4.317 vs. 4.383; it also beats 15M+-parameter PatchTST on all three ETT sets (0.781, 1.467, 0.488). The authors frame this as scale not winning everywhere, and report an efficiency ratio: v0.5 is about 26 times more parameter-efficient than TimesFM and about 2 times more efficient than PatchTST. Every number is produced under one protocol with the same splits, windows, and MASE denominator; PatchTST trains per dataset with the official code and hyperparameters, and TimesFM uses its public 200M-parameter checkpoint. The authors note Chronos-T5-large and Timer were discussed only qualitatively because their inference was intractable on their hardware.
The deployment path is complete: training takes about 12 hours on one NVIDIA T4 (Google Colab) (11.7 h for v0.3, 12.2 h for v0.5), inference needs no GPU, and on an Apple M4 CPU PyTorch FP32 full inference runs at 19.5 ms, ONNX Runtime FP32 at 10.7 ms, ONNX Runtime INT8 at 33.3 ms, and a streaming update at 19.1 ms; exports are 27.9 MB FP32 and 9.2 MB INT8, with Docker and stateful streaming inference provided. The DeltaNet layer's matrix state persists across predict() calls, so a streaming update costs one forward pass instead of reprocessing the full context for every new forecast as window-based methods must. Latency figures are measured in Table 3 on an Apple M4 CPU; code, pretrained checkpoints, and the evaluation framework are released under Apache 2.0.
Perspective
The result targets forecasting settings that must train on consumer hardware and infer on CPUs or edge devices, such as edge boxes, live analytics, and embedded boards; the authors argue that where size, cost, and deployability matter, a properly trained small model can stand in for a 200M-parameter server model. The applicable setting is univariate forecasting with a 512-timestep context and forecast horizon 96, evaluated on six public datasets. The authors' next steps include scaling the shared trunk (more channels, longer context) with the same corrected pipeline to close the high-cardinality gap while staying under 20M parameters; adding baselines such as Chronos-T5-large, Moirai, and Lag-Llama under the identical protocol; fine-tuning for other horizons and evaluating streaming mode under drift; and benchmarking the exported ONNX models on Raspberry Pi 4-class devices and microcontrollers to quantify the streaming deployment envelope.
Quantile calibration remains an open question to watch: under the same protocol v0.5's intervals run narrow, with the nominal 80% p10-p90 band covering 51.3% of held-out values versus 90.3% for v0.3, so the authors advise reading v0.5 quantiles as relative uncertainty rather than calibrated probabilities and leave recalibration or conformal post-processing to future work, noting point forecasts are unaffected. Other things to watch: overall MASE of 1.704 still trails the large models tested (TimesFM 1.447, PatchTST 1.554); the model trails clearly on the high-cardinality electricity (2.029 vs. 0.923) and traffic (1.805 vs. 0.765) sets; the fixed 512-timestep context may limit very-long-range dependencies; the model treats each channel independently and does not model cross-channel dependencies; evaluation covers only six public datasets, so results in finance, healthcare, or climate may differ; and one checkpoint per configuration is released with no seed-to-seed variance reported. In addition, the loaded text has missing entries in the dataset list and the deployment pipeline section; if those missing entries contain additional datasets or deployment details, the coverage of the related conclusions should be checked against the original.
