Google's 400M-parameter TabFM tops all 51 TabArena datasets zero-shot, and TabFM-Auto adds Elo with frozen weights
Synopsis
TabFM is a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning, is pretrained entirely on synthetic tables generated from structural causal models, and produces calibrated zero-shot predictions in a single forward pass; across all 51 TabArena datasets (38 classification, 13 regression) zero-shot TabFM ranks first among default tabular foundation models and outperforms tuned AutoML pipelines, while its two extensions, TabFM+ (multi-view feature expansion, ensembling, and post-hoc calibration) and TabFM-Auto (LLM-guided, dataset-specific data processing and feature engineering), improve both tracks over the same frozen weights.
Interpretation
TabFM recasts tabular prediction as in-context learning in a single forward pass: at the input layer dyadic feature grouping pairs each column with neighbors at offsets and projects cell values through learned Fourier frequency banks, then linear-time column-wise Induced Self-Attention Blocks alternate with row-wise Self-Attention Blocks using RoPE, eight learned CLS tokens pool arbitrary feature counts into fixed-width row representations, and a 24-layer, width-1024 in-context predictor produces the output. Earlier in-context tabular prediction was constrained by naive attention over rows and columns, which historically confined tabular transformers to fewer than 1,000 rows; by decoupling feature encoding from cross-instance reasoning, TabFM scales context to 16,384 instances, an order of magnitude beyond the regime in which this paradigm was first demonstrated. The paper gives a full architecture specification table: spectral projection 0.05M, column embedder ISAB 3.3M, row interaction SAB 1.6M, ICL predictor 402.7M, prediction head 1.05M, totaling 408.7M parameters; it also states that an asymmetric mask admits only the labeled prefix as keys, making test predictions conditionally independent given the context and preventing label leakage.
TabFM is pretrained entirely on synthetic tables drawn from structural causal models: each dataset samples a directed acyclic graph with randomized functional dependencies, root variables propagate through non-linear transformations and algebraic aggregations to produce tables mixing continuous and discrete columns, and each sample jointly randomizes table shape, the share and cardinality of categorical columns, missing entries, label noise, and class balance, with the feature axis capped at 100 columns; a four-stage curriculum grows context from 2,048 to 16,384 instances, doubling rows per table and halving batch size at each transition to hold tokens per step fixed. This continues the Prior-Data Fitted Networks line of approximating Bayesian posterior predictives with synthetic priors, but randomizes jointly with the causal graph so evaluation tables stay inside the support of the pretraining distribution, and aligns the pretraining objective with the inference interface: every sampled table is split into a labeled context and a query block before it reaches the model, and the loss is evaluated only on query rows. The paper states explicitly that training data are all SCM-generated synthetic tables and gives the row-count and batch-size relationship across curriculum stages; this is method description and design argument rather than independent ablation evidence.
On TabArena's 51 datasets (38 classification, 13 regression), zero-shot TabFM leads default foundation models on both tracks: regression reaches 2055.2 Elo, ahead of EXAONE-Tabular at 1973.1, TabPFN-3 at 1866.6, AutoGluon 1.5 extreme at 1851.2, and TabICLv2 at 1723.6, with the lowest geometric-mean error at 16.40; classification reaches 1768.6 Elo, ahead of EXAONE-Tabular at 1768.0, AutoGluon 1.5 extreme at 1669.7, TabPFN-3 at 1641.9, and TabICLv2 at 1586.4, again with the lowest geometric-mean error at 0.0957, the most outright wins at 5.54, and the lowest oracle improvability at 6.10%. Pooled over both suites TabFM ranks first among default foundation models at 1785.4 Elo, winning 65.7% of head-to-head folds against EXAONE-Tabular, 72.1% against AutoGluon 1.5 extreme, 75.9% against TabPFN-3, and 79.8% against TabICLv2. Relative to prior tabular foundation models, the lead is larger on regression, which the paper attributes to spectral embeddings placing targets on a continuous scale rather than on piecewise-constant splits; on classification the top four methods fall within a roughly 130-Elo span, and TabFM's advantage comes from lower regret across datasets, reducing oracle improvability from 9.47% for EXAONE-Tabular and 10.00% for AutoGluon 1.5 extreme to 6.10%. Evaluation uses TabArena's repeated 10-fold cross-validation, with a comparison pool of 67 method configurations (tuned tree ensembles, tuned deep baselines, four-hour-budget AutoML, and published tabular foundation models); ratings come from a Bradley–Terry fit at the level of a single (dataset, fold) pair, weighting every dataset equally, anchoring Random Forest at 1000, and taking the median rating over 100 bootstrap rounds.
Two extensions keep improving performance over frozen weights: TabFM+ standardizes and cleans columns, builds a multiplicative cross-feature pool and a truncated-SVD structural feature pool, spends half its member budget on unaugmented original columns and half on appended random draws from those pools, applies an independent view transformation per member (alternative preconditioning, random column permutations, random bijective categorical index permutations, cyclic shifts of classification labels), and stacks members with Non-Negative Least Squares plus Platt calibration for classification; TabFM-Auto pairs the frozen TabFM with Gemini-3.8-Flash in a closed program-synthesis loop that iteratively writes, evaluates, and refines a Python pipeline around the pretrained model, starting every run from the identity program (vanilla TabFM) and stopping after 96 evaluations or six hours. TabFM+ exploits the model's sensitivity to column order, numerical scaling, and appended feature interactions to produce distinct internal views from one set of weights; TabFM-Auto attributes improvements primarily to feature construction rather than hyperparameter tuning, for example Strouhal and Reynolds numbers on airfoil_self_noise or ICD-9 diagnoses folded into chapters on Diabetes130US, with median error reductions of 2.14% when the final program changed the table and 0.27% when it did not. TabFM-Auto, TabFM+, and TabFM take the top three classification positions at 1940.7, 1838.0, and 1768.6 Elo, and the top three regression positions at 2392.4, 2189.2, and 2055.2; TabFM-Auto improves 41 of 51 datasets, beats inference-time ensembling on 40 of them, improves all 13 regression datasets and reaches 0.00% oracle improvability, while on the ten classification datasets that do not improve over zero-shot TabFM test error rises by at most 1.8%.
Perspective
This work is aimed at users who need usable tabular predictions without task-specific tuning: zero-shot TabFM produces calibrated predictions in a single forward pass and suits use as a default baseline; TabFM+ suits settings willing to pay inference-time ensembling cost to extract more accuracy from the same frozen weights; TabFM-Auto suits settings that allow an LLM to synthesize dataset-specific feature-engineering pipelines inside a closed cross-validation loop, with candidate programs executed in a sandbox without access to held-out test folds. The evaluation scope is TabArena's 51 datasets (38 classification, 13 regression), and pretraining covers synthetic numerical and categorical tables of at most 16,384 rows and 100 columns. The paper's stated next steps include scaling pretraining to million-row and thousand-column tables with hierarchical row pooling and sparse feature attention, combining synthetic causal priors with real-world tabular corpora and pretrained text and temporal encoders to support free-text fields, timestamps, and high-cardinality identifiers, and generalizing the architecture from single flat tables to relational schemas so databases need not be manually joined and flattened before inference.
Pretraining covers only synthetic numerical and categorical tables of at most 16,384 rows and 100 columns, larger tables rely on length generalization or subsampling, and free-text fields still lack semantic tokenization, all of which the paper itself lists as open directions. Full details of TabFM-Auto are referred to another paper (Fu et al. 2026a), and this text only summarizes the method, so its search process and failure modes cannot be fully checked within this scope. On classification the top four methods fall within roughly a 130-Elo span, and the zero-shot advantage shows up mainly as lower oracle improvability, so readers may watch whether that gap holds across other benchmarks and data distributions. The gains of TabFM+ and TabFM-Auto come from inference-time ensembling and LLM-guided feature engineering respectively, and their cost-benefit tradeoff needs to be judged against a specific deployment budget.
