π² builds reasoning data from Wikipedia tables, lifting gpt-oss-20b and Qwen3-4B long-context reasoning by +6.25% and +3.37% average absolute accuracy
Synopsis
The work introduces π², a QA curation pipeline that extracts and expands tables from Wikipedia, generates complex reasoning questions from those tables plus relevant metadata with answers automatically determined and validated through dual-path code execution, and back-translates chain-of-thought solutions grounded in realistic context; supervised fine-tuning with gpt-oss-20b and Qwen3-4B-Instruct-2507 on π² yields consistent improvements across four long-context reasoning benchmarks and the alike π²-Bench, with average absolute accuracy gains of +6.25% and +3.37% respectively, and analysis indicates reasoning style contributes little while faithful reasoning patterns from back translation and grounded realistic long context are crucial.
Figure 1: Tables bridge realistic documents and complex reasoning in open-domain tasks. They represent data across sources , while being compatible with programmatic operations .
arXivInterpretation
It proposes π², a QA curation pipeline for long-context complex reasoning that comprises extracting and expanding tables from Wikipedia, generating complex reasoning questions from the collected tables together with relevant metadata, and back-translating chain-of-thought solutions grounded in realistic context. Compared with prior practice, the pipeline anchors data in structured realistic context drawn from Wikipedia tables and their metadata rather than generic text or hand-written items. The abstract lays out the three stages as numbered steps and states that code, data, and models are fully open-source.
Question answers are automatically determined and validated through dual-path code execution, giving generated complex reasoning questions verifiable answers at construction time. By routing answer correctness through two code execution paths, large-scale automatically built reasoning QA gains a verification mechanism. The abstract states answers are "automatically determined and validated through dual-path code execution".
Supervised fine-tuning on π² yields consistent improvements for gpt-oss-20b and Qwen3-4B-Instruct-2507 across four long-context reasoning benchmarks and π²-Bench, with average absolute accuracy gains of +6.25% and +3.37% respectively. It reports directionally consistent quantitative gains on two models of different scale, indicating the benefit is not a one-model or one-benchmark artifact. The abstract reports consistent improvements across four benchmarks plus the alike π²-Bench and gives two specific average absolute gain values.
Deeper analysis indicates reasoning style contributes little, whereas faithful reasoning patterns discovered by back translation and grounded realistic long context are crucial for the improvement. It shifts the attributed source of gains from surface reasoning style toward faithful reasoning patterns and realistic context grounding in the data, offering a testable focus for future data design. The abstract states this as an observation ("we observe") based on the authors' own experiments.
Perspective
The work targets question-answering settings that require long-context complex reasoning, especially question types whose answers can be supported by tables and metadata and verified by code execution; it suits research and engineering teams aiming to improve long-context reasoning through supervised fine-tuning, provided they can obtain structured realistic context similar to Wikipedia tables. The abstract states that code, data, and models are fully open-source, so the pipeline can be reused and extended under the same data-construction paradigm.
The visible text is abstract-level information: the specific names of the four long-context reasoning benchmarks, the composition of π²-Bench, the concrete implementation and validation criteria of dual-path code execution, and the comparison setup behind the conclusion that reasoning style contributes little are not expanded in the abstract; these are the questions a reader would continue to watch when assessing the scope of the conclusions.
