Skip to main content
Back to timeline
arXivSource publication:

Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis

Synopsis

This work first shows through controlled experiments on PRMBench and ROSCOE that supplying the correct answer does not consistently improve reasoning quality, locating the flaws in the structure of reasoning traces rather than in answer awareness, and then proposes CRAFT: it rolls out K candidate traces per sample, extracts consensus terms via TF-IRF and filters steps by z-score, converts each trace into a Reasoning Knowledge Graph with steps as nodes and logical relations as edges, aggregates them into a consensus graph, and synthesizes one new trace by traversing that graph in topological order; across FLD, FOLIO, GSM8K and OlympiadBench with GPT-5.

Source-provided article image: Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis
Figure 1 ·

Figure 1: Problem background. LLMs’ correct label-prediction is related to, but not fully determined by, whether reasoning steps are correct. Case studies of logical and mathematical reasoning are in Appendix M .

arXiv

Interpretation

Providing the correct answer does not reliably improve reasoning: on PRMBench, Simplicity and Soundness show no significant difference across the four models, only Sensitivity reaches marginal significance for o4-mini (p=0.028) and DeepSeek-R1 (p=0.012) with opposing signs; on ROSCOE only eSNLI improves consistently for all four models (p<0.01, d=0.23–0.31), DROP shows no significant influence, and some settings even degrade (e.g., GPT-5.4-nano Soundness F1 drops from 0.79 to 0.72). Where a correct answer is conventionally read as evidence of correct reasoning, this work turns that assumption into a testable controlled comparison across two complementary benchmarks and four backbone models, concluding that the flaws are structural rather than a matter of answer awareness. Two benchmarks (three PRMBench error categories, four ROSCOE datasets), four backbone models, identical prompt templates, paired Wilcoxon signed-rank tests with 95% bootstrap confidence intervals; the eSNLI improvement is attributed to NLI labels providing genuine help, with models biased toward outputting Neutral in the w/o answer setting.

CRAFT applies cross-trace consensus to both content and structure: Module I extracts consensus terms via TF-IRF and removes outlier steps by z-score (default γ=−1.0); Module II converts each trace into a Reasoning Knowledge Graph, weights edges as an equal-weighted sum of LLM confidence and Jaccard similarity between consecutive steps (λ=0.5), filters low-weight edges and isolated nodes, and aggregates by set union; Module III generates one new trace step by step in topological order. Prior remedies either target a single domain (e.g., logical reasoning only) or assume one flaw type holds for all samples (handling only redundancy or only missing steps), whereas this framework addresses Step Internal Flaws and Step-wise Flaws in one training-free pipeline. The algorithm and hyperparameters (K=5, α=0.01, β=0.3, γ=−1.0, λ=0.5, θ=0.3) are stated explicitly; filtering statistics show z-score filtering removes 6–8% of steps on logical benchmarks and 35–54% on mathematical ones, RKG structural filtering adds 20–26% removal on logical benchmarks (71–80% of all deletions), and edge filtering accounts for 57–81% of structural removals across benchmarks.

On label-prediction accuracy CRAFT is best or second-best across four datasets and two backbones: with GPT-5.4-nano, FLD 71.6, FOLIO 89.6, GSM8K 96.0, OlympiadBench 73.8; with o4-mini, FLD 75.6, FOLIO 88.8, GSM8K 98.0, OlympiadBench 73.2; on OlympiadBench it exceeds the next-best baseline by up to 8.5% accuracy while using over 70% fewer steps. Against thirteen training-free baselines spanning voting/selection, iterative refinement, search, symbolic decomposition and structural/causal trace analysis, CRAFT synthesizes a new trace rather than selecting among candidates, yielding higher accuracy together with fewer steps. 500 samples per dataset with random seed 42, evaluation by Exact-Match with LLM-As-A-Judge fallback and 200 manually verified cases; 95% Wilson confidence intervals are reported, with non-overlapping intervals against most baselines on OlympiadBench (GPT-5.4-nano), FOLIO and FLD, while easier benchmarks such as GSM8K show ceiling effects and partial overlap.

Synthesized traces score higher under fine-grained evaluation: ROSCOE shows gains in grammar (up to +5.9%) and reductions in step-level and word-level redundancy (roughly −1.2% to −2.8%); FineLogic and ReCEval compare synthesized traces against raw reasoning on validity, relevance, atomicity, entailment and non-contradiction. The improvement is extended from the final answer to intermediate step quality, directly addressing the starting observation that a correct answer does not imply correct reasoning. ROSCOE and ReCEval are run on GPT-5.4-nano and o4-mini, FineLogic on GPT-5.4-nano, with Gemini-3.1-Flash-Lite as LLM judge; ablations on OlympiadBench show drops of 20% or more when the whole pipeline is removed, −6.2%/−17.2% without RKG, −21.8%/−15.4% without synthesis, and only about −2% when Jaccard is replaced by embedding cosine similarity.

Perspective

The result targets sequential Chain-of-Thought settings and applies to logical reasoning (FLD, FOLIO) and mathematical reasoning (GSM8K, OlympiadBench) tasks with explicit labels or unique answers, validated on two backbones, GPT-5.4-nano and o4-mini; the pipeline is training-free and costs about 18 API calls per sample (K=5 plus RKG construction and stepwise synthesis), comparable to multi-sample baselines at K=10. For readers aiming to improve trace quality rather than only answer accuracy, it provides directly reusable filtering and synthesis mechanisms; for parallel or tree-structured reasoning, or for ternary-label settings (FOLIO's UNKNOWN examples are excluded), applicability needs separate confirmation.

Consensus is not correctness: the authors note that when many candidates share the same wrong reasoning pattern, consensus can amplify rather than suppress errors, and increasing K mitigates but does not eliminate this; TF is computed over only K traces per sample, so with small K the frequency signal may lack discriminative power and term-importance estimates become noisier; the K-sensitivity analysis shows FLD accuracy rising gradually from K=3 while FOLIO peaks around K=3 and then fluctuates near its highest value. The paper also notes that the API used was deployed by a collaborating company on its own GPU clusters, making performance very different from publicly available versions, which matters for reproduction and cross-comparison; FOLIO's UNKNOWN examples are excluded, so behavior under ternary logic remains an open question.

Sources