Skip to main content
Back to timeline
arXivSource publication:

TabSieve's select-then-predict framework lifts classification by 2.92% and regression by 4.45% on average across 75 classification and 52 regression tables

Synopsis

The authors propose TabSieve, a select-then-predict framework in which, given a table and a query row, the model first selects a small set of informative rows as evidence and then predicts the missing target conditioned on that evidence; to enable this capability they build TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables with a strong teacher model and strict filtering, and introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards and stabilizes mixed regression and classification training via dynamic task-advantage balancing; on a held-out benchmark of 75 classification and 52 regression tables, TabSieve consistently improves performance across shot budgets, with

Source-provided article image: TabSieve: Explicit In-Table Evidence Selection for Tabular Prediction
Figure 1 ·

Figure 1: We contrast three prediction paradigms. Traditional models cannot explicitly interpret how context is used. Strong LLMs with in-context prompting can be distracted by noisy context . TabSieve first performs evidence selection and then conducts noise-filtered reasoning to produce the final prediction .

arXiv

Interpretation

TabSieve makes evidence usage in tabular prediction explicit: it first selects a small set of informative rows as evidence, then predicts the missing target conditioned on that evidence, making the process auditable. Existing tabular models typically perform instance-wise inference and LLM-based prompting is often brittle, with models not consistently leveraging relevant rows and noisy context degrading performance; TabSieve separates evidence selection and prediction into two explicit steps. The paper describes the method as a select-then-predict framework and validates it on a held-out benchmark; the abstract does not give algorithmic details of the selection step.

The authors construct TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables using a strong teacher model with strict filtering. This dataset supplies supervision for the select-then-predict capability, letting the model learn explicit evidence-selection behavior. The data come from 331 real tables, synthesized with a strong teacher model under strict filtering; the abstract does not disclose the filtering criteria or the composition of the trajectories.

The authors introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards and stabilizes mixed regression and classification training via dynamic task-advantage balancing. Compared with optimizing only final prediction correctness, this recipe also makes evidence selection an optimization target and addresses instability in mixed regression and classification training. The abstract describes the separate-reward design and dynamic task-advantage balancing but gives no ablation numbers or reward weights.

On a held-out benchmark of 75 classification and 52 regression tables, TabSieve consistently improves performance across shot budgets, with average gains of 2.92% on classification and 4.45% on regression over the second-best baseline; further analysis indicates it concentrates more attention on the selected evidence, improving robustness to noisy context. The gains hold across shot budgets, indicating the benefit does not depend on a single context size, and the attention analysis links the improvement to evidence actually being used. The benchmark comprises 75 classification and 52 regression tables, and the reported figures are average gains over the second-best baseline; the abstract gives no variance, significance tests, or per-table results.

Perspective

The work targets settings where a table is the input and a missing target must be predicted, covering both classification and regression and evaluated across shot budgets. It makes evidence usage explicit and auditable, so it is most valuable to readers who need to explain which rows a model relied on; separating evidence selection from prediction correctness also offers a reusable recipe for training tabular models on mixed task types.

The abstract does not disclose the specific evidence-selection algorithm, the filtering criteria and trajectory composition of TabSieve-SFT-40K, the reward weights and ablations for TAB-GRPO, or per-table results, variance, and significance tests; the exact measure used in the attention analysis is also not described. These are the questions a reader would still watch for when reading the full paper.

Sources