Skip to main content
Back to timeline
arXivSource publication:

Tabular foundation models can drop 90% of context: one linear layer pulls activations back toward the full-context teacher

Lead

When a tabular foundation model's context is cut to a tenth, a single linear transformation trained on synthetic unlabeled data maps its intermediate activations toward those of a full-context teacher, beating the unaligned student on most conditions across 38 classification datasets and recovering nearly half the teacher's predictive advantage.

Source-provided article image: Closing the Context Gap: Activation Alignment for Tabular In-Context Learning
arxiv.org

Story

A tabular foundation model's predictive quality falls with the number of context examples, and activation alignment lets a student that sees only a subset move its internal representations toward a teacher that sees all of them, so inference runs on fewer context examples with better predictions. Previously, cutting this inference cost meant either compressing the context or pruning redundant transformer layers, and both routes demand substantial GPU training. Across 38 classification datasets from TabArena, evaluated on TabPFN-3 and TabFM, the aligned student outperforms the unaligned baseline in 81% of conditions with statistically significant gains.

The aligner is a linear module that predicts the residual between student and teacher activations rather than the teacher activation itself, with weights initialized near zero so the student's original representations are preserved at the start of training. Feature-based distillation usually treats the learned regressor as training-time scaffolding discarded once student weights update, whereas the aligner leaves all weights frozen and is itself the deployed inference-time intervention. Training uses only synthetic unlabeled queries and converges in seconds to minutes on commodity CPU cores without a GPU; alignment is applied at the final layer of 24-layer models, with a separate aligner per estimator for TabPFN's 8-estimator ensemble and a single estimator for TabFM.

What to watch

A next step is to test whether the same aligner transfers directly to regression, since it trains only on unlabeled synthetic queries and carries no classification-specific assumption. The aligner could also be attached to models compressed by context compression or layer pruning, pulling representations perturbed by compression back toward the full-capacity teacher while keeping the inference speedup. For a student whose context is chosen by retrieval, the same approach can correct representations without changing the input.

Five of the 38 datasets show a negative mean effective-sample gain for both models, and there is currently no characterization that separates them from the datasets that benefit. Training the aligner requires offline access to the full dataset to extract teacher activations, so it addresses inference-time compute and memory bottlenecks rather than genuinely small datasets. The method has been validated only on classification tasks and two tabular foundation models, leaving regression and other architectures as open questions.

Sources