Skip to main content
Back to timeline
arXivSource publication:

How Test-Time Compute Lifts Tabular Foundation Models: DiagScale Adaptation and Greedy Aggregation Give Consistent Gains, While Context Construction Depends on the Task

Synopsis

This work systematically studies how test-time compute affects predictions of pretrained tabular foundation models along three axes—adaptation, aggregation, and context construction: it introduces DiagScale, a diagonal query-key similarity update training only 0.003-0.03% of parameters that matches full fine-tuning gains across three independently pretrained backbones; on a broader pool of 96 configurations, greedy selection reduces error by 2.4% relative to the default predictor while uniform averaging increases error; attention-guided retrieval improves TabPFN-3 on some large tables, whereas the context expansion methods tested yield no consistent improvement.

Source-provided article image: Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits
Figure 1 ·

Figure 1: Three families of test-time compute modify different components of the inference pipeline. Top: an inference pipeline comprising preprocessing, a frozen tabular foundation model, and combination of predictions across data views. Bottom: context construction changes the conditioning data; adaptation changes model parameters; aggregation changes the prediction pool and the rule used to select or combine its members.

arXiv

Interpretation

Introduces DiagScale, an adaptation method in the form of a diagonal query-key similarity update that trains only 0.003-0.03% of model parameters yet achieves gains comparable to full fine-tuning across three independently pretrained backbones. Relative to full fine-tuning, it compresses trainable parameters to a very small fraction while keeping comparable performance, indicating adaptation gains need not rely on large-scale parameter updates. Evaluated on modern tabular foundation models across the TabArena benchmark, supplemented by experiments on wide and large-scale tables from OpenML; gains reproduced across three independently pretrained backbones.

Aggregation depends on both pool composition and selection strategy: TabPFN-3 already averages predictions from different preprocessing variants of the same data, and adding more such predictions yields diminishing returns; with a broader pool of 96 configurations, greedy selection reduces error by 2.4% relative to the default predictor, but uniform averaging increases error. Decomposes aggregation into pool composition and selection strategy, showing a directional difference—greedy selection helps while uniform averaging hurts. Systematic evaluation on the TabArena benchmark with OpenML wide and large-scale table experiments; reports the error change relative to the default predictor.

For context construction, attention-guided retrieval improves TabPFN-3's predictions on some large tables and supports source pools beyond the full context memory limit; the context expansion methods tested yield no consistent improvement. Distinguishes retrieval-based context construction from context expansion, showing the former helps in specific large-table settings and can exceed memory limits, while the latter lacks consistent benefit. Experiments on TabArena and OpenML large tables; effects vary with task and data regime, with no consistent cross-setting gain reported.

Adaptation and aggregation over the same backbone yield further gains when combined, but require substantially more computation than default inference; these trade-offs motivate choosing strategies according to the available computation budget. Places the gains and computational costs of the three test-time compute strategies in one framework, offering practical guidance for budget-based trade-offs. Synthesizes benchmark-level results, indicating adaptation and selective aggregation yield consistent benchmark-level gains while context construction benefits depend more on task and data regime.

Perspective

The work targets researchers and practitioners using strong pretrained tabular foundation models, in evaluation settings centered on the TabArena benchmark and supplemented by OpenML wide and large-scale tables. Its conclusions are organized around available compute budget: adaptation and selective aggregation yield consistent benchmark-level gains, and combining them yields further gains, but at substantially higher computation than default inference; context construction benefits depend more on the specific task and data regime. Attention-guided retrieval helps on some large tables and supports source pools beyond the full context memory limit, making it relevant for memory-constrained large-table settings.

Several open questions remain: under which tasks and data regimes context construction benefits appear reliably still needs finer characterization; whether the 2.4% error reduction from greedy selection on the 96-configuration pool changes with pool size and configuration diversity is not elaborated in the text; and at what budget the extra computation from combining adaptation and aggregation remains worthwhile is left to deployment-specific trade-offs. In addition, this reading is at summary scope and does not include figures or full experimental details, so the specific behavior of each strategy across data regimes still requires consulting the original.

Sources