Skip to main content
Back to timeline
arXivSource publication:

Systematic evaluation of streaming tabular foundation models: TabFM+FIFO leads on nearly all nine streams but serves over an order of magnitude slower than the fastest ensemble

Related research and updates

Synopsis

In a unified prequential (test-then-train) protocol across real and synthetic streams, this study systematically compares tabular foundation models (TFMs) with streaming learners and streaming AutoML, finding that TFMs reach the highest predictive accuracy, recover faster after drift, and stay most accurate under label delay, but cost more than an order of magnitude more to serve; memory size and backbone quality matter more than the memory management policy, and repeated context encoding is the dominant cost, which a compact causal model with cached encoded context can hold nearly constant across memory sizes at close accuracy.

Source-provided article image: Rethinking Tabular Foundation Models On Data Streams
Figure 2 ·

Figure 2: Accuracy over the synthetic streams, each with its own drift type, in stages of roughly a tenth of the stream. Dashes mark concept changes.

arXiv

Interpretation

Under the same prequential protocol, TFMs achieve the highest predictive performance: TabFM+FIFO ranks first in accuracy on every stream but the abrupt-drift one, with mean accuracy .869 and mean rank 1.11, while Drift-Resilient TabPFN is narrowly ahead on the abrupt-drift stream. Prior evaluations of DualFIFO and CURE neither compared across drift types nor placed serving cost beside classical learners; this study contrasts backbone, memory policy, drift type, and cost under one protocol and one memory budget. Per-stream accuracy tables over nine streams (five real, four synthetic covering abrupt, gradual, incremental, and recurring drift), plus two-sided exact Wilcoxon signed-rank tests with Holm correction on paired per-stream accuracies; the authors note that with few streams the test does not separate the leading in-context methods, so they emphasize the descriptive ranking.

Memory size affects predictive performance substantially more than the choice of memory management policy: replacing CURE or DualFIFO with FIFO on the same backbone changes accuracy little at any budget while reducing serving cost; increasing the budget improves accuracy more noticeably but also raises latency, and the abrupt-drift stream peaks below the largest budget. Prior work invested in more elaborate memory management (class-balanced long-term store, redundancy-based eviction); this controlled one-variable sweep shows limited marginal return from those policies and favors tuning memory size. Comparisons of FIFO against CURE and DualFIFO on TabICL v2 and TabPFN v1 for noaa and agr_a, reporting accuracy and milliseconds per instance, plus a sweep over 100, 250, 500, and 1,000 retained rows.

Under label delay, TFMs retain the highest mean performance, but the advantage comes from starting higher rather than degrading less; CURE becomes beneficial only on the smart-meter stream, where it increasingly outperforms FIFO on the same backbone and surpasses TabFM+FIFO at the longest delay. Prior TFM stream methods assumed immediate labels; this study treats delays of 0, 100, and 1,000 as an independent variable and separates the starting-point explanation from the degradation-rate explanation. Per-stream accuracy tables for four streams (noaa, meter, agr_a, tree_r) at three delays; the authors state that the CURE benefit is confined to one stream and do not claim that delayed feedback generally favors elaborate memory management.

Repeated context encoding is the dominant serving cost of streaming TFMs: a fixed per-call overhead dominates single-instance serving time and is largely context encoding; with a compact model using a causal datapoint axis, cached serving stays flat in time and GPU memory as memory grows, whereas re-encoding grows superlinearly in time and close to quadratically in memory, taking up to an order of magnitude longer at the largest budget. Existing TFM backbones attend bidirectionally over rows and cannot incrementally reuse encoded rows; this study brings causal attention and KV-cache ideas from language models into a compact TFM as a proof of concept that the bottleneck is removable. A fit on noaa of fixed overhead and per-query cost (TabICL v2 fixed 26.1 ms, TabPFN v1 59.4 ms, TabFM 120 ms), plus accuracy, time, and peak GPU memory for cached versus re-encoded serving of one checkpoint on four streams; the authors note the models are small and the cache unoptimized, so the saving is far below the theoretical ceiling.

Perspective

The study targets streaming tabular classification, in settings where data arrive in arrival order and a model must predict at any time while adapting under bounded memory and time, such as fraud and intrusion detection, sensor monitoring, and clinical decision support. For practitioners, the directly usable conclusions are to use FIFO on a strong backbone, tune memory size first, and use query batching to amortize encoding cost where latency allows; for researchers, it frames streaming TFM efficiency as an architectural problem and offers cached reuse as a direction to pursue. Synthetic streams pair each drift type with a different generator, so they describe stream conditions rather than supporting direct comparisons between drift types.

Per-stream differences are large: gradual and incremental streams separate almost nobody, and under recurring drift the swing within a single method exceeds the gap between paradigms, so aggregate rankings should be read alongside temporal behavior. Because the number of streams is small, the statistical test does not separate the leading in-context methods, leaving the leading result descriptive. The caching benefit comes from compact models the authors pretrain themselves, and whether existing large TFMs support similar reuse remains open; cached rows may retain information from observations that have since been evicted, so cached and re-encoded predictions need not agree exactly. Unlabelled observations arriving during label delay are not used, and substantially larger memories and other tabular tasks such as regression remain unexplored.

Sources