Retro makes tabular foundation models refine predictions earlier and more broadly, reaching top-three overall Elo on TabArena, TALENT, and RelArena
Synopsis
Tracing per-query predictive changes through TabICLv2, TabPFN-3, and EXAONE-Tabular, the authors find refinement is highly uneven across depth and often concentrated in later layers; they propose retrospective inference and build Retro, which uses Attention Residuals to adaptively reweight contributions from different depths and query-conditioned gated attention to modulate contextual updates element-wise, shifting refinement earlier and more broadly and reaching top-three overall Elo on TabArena, TALENT, and RelArena while lying on the performance-cost Pareto frontier.
Figure 1: Retro lies on the Pareto frontier of three mainstream benchmarks.
arXivInterpretation
The work characterizes sample-level predictive dynamics in strong tabular foundation models: using a linear readout fixed at the final layer to track each query's true-class probability margin, it finds refinement is highly uneven, with the final third of depth accounting for 84.4% of absolute margin movement in TabICLv2, versus 51.5% for TabPFN-3 and 61.1% for EXAONE-Tabular. Prior analyses probed representations or layers to reveal redundancy and non-uniform contributions; this work instead traces how individual queries change predictive state, making 'which queries are revised when' the design motivation. Run on the first official split of 38 TabArena classification datasets for three models, using the same cross-fitted support representations and a frozen readout; statistics include all query rows with equal-weight averaging across datasets, and visual cases are illustrative.
It proposes retrospective inference and instantiates it in Retro: Attention Residuals let later stages adaptively reweight, per row, the initial encoding and sums of completed two-layer update groups, while query-conditioned gated attention modulates attention outputs element-wise before projection, yielding earlier and more distributed predictive refinement. In a standard residual stack an early contribution can only influence later prediction as part of an accumulated state and cannot be selected independently; Retro keeps intermediate contributions as separately weightable sources and additionally controls how new contextual information rewrites each query. Under the same frozen readout, the ungated retrospective variant shows mean correctness turnover per transition of 24.66% (9.82% for TabICLv2), 79.5% of queries changing correctness at least once (63.8% for TabICLv2), and an early-motion share of 60.7% (15.6% for TabICLv2), larger on 36 of 38 datasets; final frozen-readout accuracy is 87.5% versus 87.3%.
Gating interventions show the trained models depend more on the direction of the contextual update than on its magnitude: replacing each query's gate with the channelwise mean gate of the unlabeled query batch reduces mean classification accuracy by 0.719 percentage points, while replacing only direction loses 0.616 points and replacing only magnitude loses 0.0011 points, a paired difference of 0.618 points (95% interval [0.173, 1.217]). This indicates gating shapes the geometry of the contextual update rather than merely scaling its overall strength, giving mechanism-level support for query-conditioned modulation. Covers the first official split of all 51 TabArena datasets (38 classification, 13 regression) using every support and query row, one view, FP32 inference, and no ensembling, with equal dataset weighting and 95% intervals from 20,000 paired dataset-bootstrap resamples; the classification effect is heterogeneous, with median 0.126 and 23 datasets favoring preserving direction, seven favoring magnitude, and eight ties.
Retro is evaluated on three benchmarks: it ranks third on TabArena, second on TALENT, and first on RelArena by overall Elo, lies on the estimated performance-cost Pareto frontier of each benchmark, and improves predictive performance over TabICLv2 with nearly unchanged inference time. Relative to its TabICLv2 backbone, Retro ranks higher in the overall rankings on TabArena and TALENT; in the ablation the combined model reaches an Elo point estimate of 1115.6, above Attention Residuals only (1093.1), Gated Attention only (969.0), and neither (822.3). TabArena includes 38 classification and 13 regression datasets (594 and 222 official splits), TALENT includes 200 classification and 100 regression datasets, and RelArena includes 12 classification and nine regression tasks; comparison pools contain 37/38/39, 17, and 10 methods, Elo values are not calibrated across benchmarks, and the ablation uses a separate comparison pool.
Perspective
The results target settings where a pretrained tabular foundation model predicts with labeled support examples as context, covering classification, regression, and relational prediction. Retro retains TabICLv2's synthetic task generator, preprocessing, and inference procedures, with separately trained classification and regression models using task-specific target embeddings and output heads, so its gains are measured under that backbone and training recipe. Retrospective access retains the initial encoding and sums of completed two-layer groups rather than every full hidden state; the aggregation modules add 24,576 parameters and gate projections add roughly 3.15M parameters, with aggregation and gate projection costs on the order of and respectively, complementing the existing support-conditioned attention and feed-forward computation. Engineers seeking better tabular prediction without materially higher inference time can borrow the two design ideas directly: keeping intermediate contributions as separately weightable sources, and modulating contextual updates element-wise.
The frozen-readout analysis measures how predictive information is expressed relative to a common final-layer decision rule rather than the total information in each representation, so it is used to compare layer-wise refinement patterns; depth indices denote native blocks, and computation per layer is not equal across models. The gating direction effect is heterogeneous across classification tasks, with median 0.126 and some datasets where direction replacement is less harmful, so its generality across more tasks and backbones remains to be observed. Gate-exchange experiments show neighborhood-constrained exchange is less disruptive, consistent with local continuity of sample-conditioned updates, but do not establish distinct semantic roles for those neighborhoods. Elo values are not calibrated across benchmarks, and TALENT intervals reflect dataset-order variability rather than evaluation-seed uncertainty or a bootstrap confidence interval, so cross-benchmark ranking comparisons should be read within each comparison pool. In addition, several equations, figure numbers, and some numeric values are not fully rendered in the parsed text, so reproducing implementation details would still require the recurrence and complete tables in the original appendices.
