Pretraining leaves a spectral fingerprint in the weights: after conditioning on architecture and strategy, one spectral statistic ranks OOD robustness across 116 foundation models without target data
Synopsis
The work reads pretrained-weight spectral structure through 15 (architecture x pretraining-strategy) cells, proves that the OOD accuracy gap is bounded by how tightly source representations concentrate, and shows that a single statistic selected per cell orders OOD robustness across 116 models and 7 modalities, with the cell, metric and sign locked before evaluation for out-of-matrix families including EEG, genomic and protein.
Figure 1: (Architecture × \times Strategy) → \to Spectral Geometry → \to OOD Robustness. Architecture and pretraining strategy jointly shape the layer-wise singular-value spectrum (2) , which determines whether representations are concentrated or diffuse (3) , setting the OOD accuracy gap (4) . The sign of ρ \rho is a property of the cell and of the statistic read in it, not of the domain: each AR-CLM cell is positive under its own calibrated metric, while BiDi-MLM/Protein, CNN and SSM-v2/RWKV are negative under theirs.
arXivInterpretation
The paper supplies a theory chain: Theorem 1 links the OOD accuracy gap to the log stable rank concentration of source features, Lemma 2 connects nested residual weights to final feature geometry, and Propositions 3-5 give architecture-conditional weight-space proxy statistics for AR-CLM, BiDi-MLM and ViT-CLIP. Earlier work tying weight spectral statistics to model quality targeted in-distribution generalization and applied one statistic to every architecture; this work formalizes the relation within (architecture x strategy) cells and extends it to OOD robustness. The proofs hold under explicit assumptions T1-T6; the authors flag the T4 concentration bridge, the HT training relation, order preservation and polarity assignment as empirically checkable assumptions rather than universal mathematical facts.
In each of the 15 primary cells, the statistic selected for that cell orders models by OOD drop, with a mean leave-one-out Spearman correlation of 0.824; across 454 within-cell pairwise comparisons the statistic identifies the more OOD-robust model in 92% of cases. Global pooling yields only 0.245, strategy-only conditioning 0.478 and architecture-only 0.745; the monotone gain shows (architecture x strategy) is the coarsest grouping that still keeps the within-cell metric-OOD relation consistent. Covers 116 models across 7 modalities; 9 of 15 cells are two-sided significant, 4 more one-sided under the pre-specified direction and 2 are marginal; leave-one-cell-out keeps the 14-cell mean between 0.800 and 0.824.
Acting on spectral concentration narrows the OOD gap: across 22 NLP models and 1,232 interventions, truncating the source-feature matrix to its top-k left singular vectors reduces OOD drop by 24% relative to the untruncated baseline while retaining 87.5% of in-distribution accuracy, and the drop rises monotonically with k. This turns the proximal variable of Theorem 1 into something directly manipulable and motivates an effect-size hierarchy in which interventions on feature concentration outperform those propagating through the longer weight-to-feature chain. Two null controls that leave feature rank unchanged, LoRA rank adaptation and rank surgery, yield no OOD benefit; metric-guided weight repair beats random targeting in 9 of 10 NLP models.
The diagnostic transfers data-free to families outside the matrix: for EEG, AMPLIFY protein and Evo2 genomic, the cell, metric and sign were locked before any OOD outcome was observed, and 3 of 3 directional predictions were confirmed; for AMPLIFY the logged rank-percentile verdict was not confirmed and only the pairwise ordering held. This constitutes a prospective test rather than post-hoc illustration, moving weight-space diagnostics from in-distribution generalization toward release-time OOD robustness triage. EEG shows the positive direction in both the BiDi-Transformer and BiMamba/linear-attention sub-groups; Evo2-7B reaches OOD NLL below random across 5 species while 1B is near chance; the BioNeMo ESM-2 outcome is pending.
Perspective
The framework is meant to rank models within a cell rather than predict absolute accuracy on a target domain, and it targets researchers and engineering teams who need to screen OOD fragility at release or model-selection time using only released weights plus architecture and pretraining-strategy labels. Architectural subtypes whose spectral scaling diverges, such as RoPE against absolute positional encoding, need their own calibration set, which the within-cell knee check detects; anchoring the direction in a domain with few public models needs a one-time calibration on reference models. The intervention side reduces OOD drop by 24% while retaining 87.5% in-distribution accuracy across 22 NLP models and 1,232 interventions, offering a starting point for matching repair method to failure mode.
A careful reader would still watch several things: the joint normalized-margin concentration in T4, the HT training relation, order preservation from weights to features and the polarity assignment in some cells are listed by the authors as empirically checkable assumptions rather than proven facts; the 15-cell p-values do not correct for selecting each cell's metric and sign on the same outcomes; the ViT-DINO and SSM-Mamba cells are marginal and their publicly available checkpoints are already exhausted; in the AMPLIFY case the rank-percentile verdict was not confirmed, indicating the calibration set must match the target model's architectural subtype; and in the two intervention cases where both were measured, feature rank rose while OOD drop fell, which the authors report as not showing mediation. In addition, many numeric values, table cells and formula symbols are missing from this parsed text, so specific correlations, confidence intervals or effect sizes should be checked against the original.
