Distilling tabular foundation models into lightweight students: 57-98 Elo above tuned-and-ensembled baselines on TabArena and up to 21.6x inference speedup
Synopsis
This work studies how to distill tabular foundation models (TFMs), whose in-context learning makes inference expensive, into lightweight dataset-specific students, and derives a recipe that uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries; on TabArena the resulting students outperform supervised tuned-and-ensembled counterparts by 57-98 Elo points, and applied unchanged to TALENT the recipe improves matched default students on 236-258 of 300 datasets while reducing median primary error by 4.0-6.4% and achieving median inference speedups of 3.0-21.6 times over their teachers.
Figure 1: Constructing supervision for TFM distillation. (Left) We study how to construct teacher supervision for ditillation (RQ1) and whether and how expanding teacher queries improves distillation (RQ2). We find that training solely on teacher predictions generated with the full labeled training set as context achieves the best performance. (Right) Elo versus median inference time per 1K samples on TabArena for default lightweight students, their distilled counterparts, and the TFM teachers, with Elo calibrated jointly over the same 76 configurations as in Section 4 .
arXivInterpretation
The paper examines two design questions for TFM distillation, namely how to construct teacher supervision and whether expanding query coverage improves distillation, studying them across two TFMs and both neural and tree-based students, and derives an effective distillation recipe. Previously, the strong predictive performance of TFMs relied on repeatedly conditioning on labeled data, making inference expensive; this work treats teacher-context construction and query-coverage expansion as explicit design dimensions to be tested rather than applying routine knowledge distillation. Evidence comes from a systematic comparison across two TFMs and two student architecture families, with the abstract reporting the derived recipe and its performance on multiple benchmarks, making this a method-level controlled study.
The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. Unlike training students on ground-truth labels or only on observed queries, this recipe bases supervision entirely on teacher predictions and explicitly adds synthetic queries to expand query coverage. The abstract states the recipe was derived from examination across two TFMs and both student types, and was applied unchanged on TabArena and TALENT.
On TabArena, the distilled students outperform their supervised trained, tuned-and-ensembled counterparts by 57-98 Elo points. The comparison baseline is a tuned-and-ensembled supervised model rather than an untuned one, so the gain is measured against a strong baseline. The abstract reports a 57-98 Elo point range, a leaderboard-style quantitative comparison.
Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets, reduces median primary error by 4.0-6.4%, and achieves median inference speedups of 3.0-21.6 times over their teachers. This indicates the recipe transfers across different TFMs and default student configurations, and quantifies the trade-off between predictive performance and repeated inference cost. Evidence consists of win counts across 300 datasets, a median error-change range, and a median speedup range, giving broad coverage.
Perspective
The results target tabular prediction settings that require repeated inference: when a model must be called many times on the same dataset and the teacher's context-conditioning cost becomes the bottleneck, distilling a TFM into a lightweight dataset-specific student is appropriate. The recipe uses the full labeled training set as teacher context and relies solely on teacher predictions for observed and synthetic queries as supervision, so it fits settings where a labeled training set exists and training a separate student per dataset is acceptable. The abstract indicates the recipe was examined across two TFMs and both neural and tree-based students and applied unchanged on TabArena and TALENT, suggesting the design is not tied to a single model family. Engineering teams seeking to cut deployment cost while retaining TFM-level predictions can consult this trade-off path directly; code is available at the link given in the abstract.
This document is abstract-level material and does not include the paper's tables, hyperparameter details, or per-dataset results, so the exact construction of synthetic queries, the effect of teacher-context size, and differences among student architectures remain to be confirmed in the full text. The abstract reports Elo point ranges, win counts, and median error changes without variance or significance information, so readers judging stability across data sizes and feature types would still need the complete experiments. In addition, distilled students are dataset-specific, and how they behave under cross-dataset reuse or cold-start scenarios is not addressed in the abstract, which is an open direction worth watching.
