TabFM-Auto pairs an LLM agent with frozen TabFM to evolve data pipelines, lifting Elo from 1785 to 2013 across all 51 TabArena datasets
Synopsis
TabFM-Auto pairs the frozen tabular foundation model TabFM with a language-model coding agent that iteratively rewrites four pipeline stages—data cleaning, feature engineering, context selection, and post-processing—guided by dataset metadata and validation feedback; across all 51 TabArena datasets five configurations take the top five overall positions, the best (Codex with Opus 5) raises TabFM from 1785.3 to 2013.0 Elo, the discovered pipelines transfer without further search to TabPFN-3, TabICLv2, and EXAONE-Tabular (+69 to +143 Elo), and TabFM-Auto ranks first overall among MLE agents on the 8 tabular competitions of MLE-Bench.
Interpretation
The paper introduces TabFM-Auto: TabFM's weights stay frozen while an LLM coding agent iteratively edits four modular Python functions (preprocess(), engineer(), sample(), postprocess()) plus the TABFM_KWARGS configuration dictionary under 3-fold cross-validation, turning column names, units, task descriptions, and auxiliary files into explicit pipeline code. Earlier LLM feature-engineering methods such as CAAFE, FeatLLM, and OCTree only append derived columns to a single table, TabFM+ only adds generic pairwise crosses, SVD projections, and multi-view ensembling, and end-to-end MLE agents such as AIDE, R&D-Agent, and MLEvolve retrain tree ensembles or neural networks during search; TabFM-Auto confines the search to the data pipeline while the predictor stays fixed. The paper gives a formal objective over the four stages, a sandbox and frozen-model verification protocol, and reports that each candidate pipeline needs no gradient training, is fast to evaluate, and carries no model-retraining noise; five agent and language-model combinations were run across 51 datasets.
On all 51 TabArena datasets (38 classification, 13 regression), five TabFM-Auto configurations take the top five overall positions, with the best (Codex with Opus 5) reaching 2013.0 Elo, above TabFM's 1785.3 Elo and TabFM+'s 1856.0 Elo, and above 4-hour AutoGluon 1.5 (extreme) at 1668.4 Elo. The paper reports larger gains on regression: TabFM-Auto raises TabFM to 2512.9 Elo and improves on all 13 regression datasets, while on classification it reaches 1966.3 Elo, with Antigravity plus Gemini 3.8 Flash achieving the lowest G-Mean test error (0.0902) and the most dataset wins (16.11). Each dataset gets one 3-fold cross-validation search on fold 0's training split (up to 96 evaluations or 6 hours on one H100 GPU), after which the pipeline is frozen and scored on all official test folds; the paper also reports fold-0-only held-out rankings and validation-to-test curves over up to 15 intermediate checkpoints.
The paper groups datasets by whether column names carry real semantics and inspects the features the agents write: on the 17 datasets whose column names or task descriptions identify physical, clinical, engineering, or economic quantities, the agents write domain formulas directly (Strouhal numbers on airfoil_self_noise, ICD-9 codes mapped to 19 organ-system chapters on Diabetes130US, Shock Index on MIC), cutting official test error by 7.25% (Opus 5) to 8.16% (Gemini 3.8 Flash), nearly three times the reduction on the 34 datasets with anonymized or generic schemas. The paper decomposes the gains into domain-knowledge features, statistical and structural features (bipartite graph degrees and co-occurrences on Amazon_employee_access, frequency encodings and missingness signals on kddcup09_appetency, collinearity pruning and truncated SVD on wide tables), and context-and-calibration-only changes, noting that 4 of the 5 runs that left the feature table untouched improved through context sampling or prior calibration. The analysis covers the final pipelines of Claude Code with Opus 5 and Antigravity with Gemini 3.8 Flash on all 51 datasets and reports mean and median error reductions per category, alongside a per-dataset official test score table.
Ablation and transfer experiments probe the mechanism: an unconstrained coding agent with the same Antigravity harness, Gemini 3.8 Flash model, and 6-hour budget that may train any models reaches only 1468.8 Elo versus TabFM-Auto's 1979.6 Elo, and running the final pipelines of the three best configurations unchanged around TabPFN-3, TabICLv2, and EXAONE-Tabular improves every model over its default, with Claude Code with Opus 5 transferring best (TabICLv2 +143.3, TabPFN-3 +130.8, EXAONE-Tabular +88.7 Elo). The paper argues from this that coding agents perform better when paired with a frozen foundation model than when training end to end, and that what transfers is a short readable program of column-level operations, sentinel rules, and target link functions rather than fitted numerical parameters. The ablation uses the same agent harness, model, and budget; the transfer experiment scores all official test folds of all 51 datasets and keeps only the context size from TABFM_KWARGS, since the other models do not support NNLS view weighting, feature crosses, and SVD features.
Perspective
This work targets researchers and engineering teams who use frozen tabular foundation models for structured-data prediction, and it applies to tabular tasks that come with metadata such as column names, units, task descriptions, or auxiliary files, as well as to tables exceeding TabFM's 16,384-row pretraining length or showing severe class imbalance. The reusable path the paper sketches is to collect formulas, cleaning rules, target transforms, and relational features across datasets into a library that warm-starts search on new tables with fewer evaluations; the authors further propose folding these operations into pretraining to obtain schema-aware tabular foundation models and extending the same search to multi-table relational databases. The transfer experiment shows the pipelines still improve other frozen tabular foundation models when only the context size is retained, so the result is directly relevant to readers who want better predictions without retraining a model.
A careful reader would still watch several things: gains shrink markedly on anonymized headers, so the method depends on readable column names and task descriptions; pipeline search requires repeated validation evaluations per dataset, and the paper reports about $17.6K in LLM API fees for five 51-dataset sweeps, so cost grows with dataset count; the paper notes that fold 0 is the only fully held-out split and that test rows of other folds overlap the training rows used for search, which could inflate cross-fold scores, with fold-0-only evaluation and intermediate-checkpoint curves used to probe this; and a few datasets show small error increases, which the paper attributes to near-saturated or small noisy tables. In addition, what was loaded here is the full paper text and the external story without the figure images themselves, so descriptions of figure details rest on the prose.
