WxFM-XL adapts univariate foundation models with band-specific error-correlation graphs and dynamic fusion, taking best MSE in all 16 multi-station weather settings
Synopsis
WxFM-XL decomposes univariate time-series foundation model forecasts into low, mid, and high frequency components, builds a cross-station error-correlation prior graph per band, dynamically fuses it with a static spatial correlation graph via learnable gates, and applies a lightweight graph residual adapter to correct forecasts over a frozen backbone; across Hunan, French MeteoNet, and global hourly station datasets, a WxFM-XL variant attains the best MSE in all 16 dataset-variable-horizon settings.
Figure 1: Visualization of error correlation dependencies for the v-component of wind speed in the French region. (a) Sundial’s forecast error correlations relative to an anchor station (star). (b) Topologies of spatial and band-specific error correlation graphs (showing top 600 edges with unique connections highlighted).
arXivInterpretation
A framework for adapting univariate time-series foundation models to multi-station weather forecasting: frequency-domain decomposition of foundation model forecasts, per-band error-correlation prior graphs, and gated dynamic fusion with a static spatial correlation graph to form frequency-aware dynamic fusion prior graphs. Prior multivariate extensions such as AdaPTS and Timer-XL treat stations as generic variables and mainly model inter-variable correlations, without explicitly modeling station spatial information or stationwise error priors relative to the foundation model; this work models both as static and dynamic priors and fuses them. The paper gives full formulations for frequency decomposition, the error-correlation graph (Pearson correlation, retaining top positive neighbors, symmetrization, unit self-loops), the spatial graph (standardized longitude/latitude/elevation, Gaussian kernel, median nonzero distance, k nearest neighbors), and gated fusion, with Appendix A.1 showing the conjugate-symmetric mask decomposition reconstructs the original forecast exactly.
A lightweight graph residual adapter: with the foundation model frozen, each band is processed by multi-scale time patch embedding, graph normalization and propagation, Time Attention plus MLP temporal refinement, patch unembedding, and sample-dependent scale-weighted fusion to output residual corrections. Unlike fine-tuning the foundation model backbone or using generic multivariate adapters, only adapter parameters are trained, preserving the foundation model's strong temporal priors while injecting station spatial information and stationwise error priors. The training objective combines final-forecast MSE with an energy-normalized RMSE per frequency band to prevent high-energy low-frequency components from dominating gradients; Appendix A.4 shows online complexity and activation memory grow linearly with the number of stations, avoiding the quadratic station dependence of dense spatial attention.
On three hourly station datasets (Hunan, French MeteoNet, global), a WxFM-XL variant achieves the best MSE in all 16 dataset-variable-horizon settings, with WxFM-XL (Timer) first in nine and WxFM-XL (Sundial) first in the remaining seven, and both variants consistently outperform their frozen backbones. Relative to task-specific models (STELLA, CDPNet, MSGNet, Corrformer) and foundation model extensions (Timer-XL, Moirai, AdaPTS), the adaptation is more consistent across regions, variables, and horizons. Results are means and sample standard deviations over three seeds (2021, 2022, 2023); Timer and Timer-XL use deterministic inference, while Sundial and Moirai show negligible variation at the reported precision so their standard deviations are omitted; global variables were scaled by 10 during preprocessing, with MSE divided by 100 and MAE by 10 before reporting.
Ablations and graph visualizations support complementary components: removing frequency-domain decomposition, the error-correlation prior graph, or the multi-scale temporal module all increase MSE; error prior graphs show heavier edge-distance tails than the spatial graph, and their Pearson correlation with the spatial graph decreases from low to high frequency. This provides direct structural evidence that a single static spatial graph cannot represent error dependencies across frequency components and that dynamic fusion is needed, beyond end-to-end metric gains alone. Ablations are run on the Hunan and French regional datasets with both MSE and MAE reported; visualizations compare edge-distance weighted ECDFs, structural correlations, and weighted edge overlap (weighted Jaccard index).
Perspective
The results target station-based hourly meteorological forecasting with fixed geographic coordinates, covering temperature and wind variables at regional (Hunan, northwestern France) and global scales, with forecast horizons of 24 and 96 hours and an input window fixed to accommodate Timer. The method uses frozen Timer or Sundial backbones and trains only the adapter, patch projections, graph gates, fusion factors, temporal modules, and scale scorers, so its gains build on the strong univariate temporal priors already present in the chosen foundation model. Error-correlation prior graphs are constructed offline from the training split, with validation and test targets excluded from graph construction, making the approach suited to settings where training-period forecast errors are available in advance. Inter-variable dependencies are left to future work; the current framework handles one variable at a time.
The error-correlation prior graphs depend on the foundation model's forecast errors on the training split, so they must be rebuilt when the backbone or forecasting setup changes, and their stability on new models or regions remains to be observed. The paper reports point-forecast metrics such as MSE and MAE, without probabilistic forecasting or extreme-event evaluation. Ablations are run on the Hunan and French regional datasets, not the global dataset. Visualization conclusions rest on graph statistics such as edge-distance distributions, structural correlations, and weighted edge overlap, and the quantitative link between these and forecast gains remains an open question. In addition, several numeric values in the main text (such as the MSE reduction on French wind settings relative to the strongest task-specific baseline, the specific MSE values before and after Timer-XL fine-tuning, and the specific graph correlation values) appear as placeholders in the provided text, so readers needing exact figures should consult the original tables and figures.
