Only content delivered through the publication boundary on this date is included.
arXiv This work first shows through controlled experiments on PRMBench and ROSCOE that supplying the correct answer does not consistently improve reasoning quality, locating the flaws in the structure of reasoning traces rather than in answer awareness, and then proposes CRAFT: it rolls out K candidate traces per sample, extracts consensus terms via TF-IRF and filters steps by z-score, converts each trace into a Reasoning Knowledge Graph with steps as nodes and logical relations as edges, aggregates them into a consensus graph, and synthesizes one new trace by traversing that graph in topological order; across FLD, FOLIO, GSM8K and OlympiadBench with GPT-5.
This work first shows through controlled experiments on PRMBench and ROSCOE that supplying the correct answer does not consistently improve reasoning quality, locating the flaws in the structure of reasoning traces rather than in answer awareness, and then proposes CRAFT: it rolls out K candidate traces per sample, extracts consensus terms via TF-IRF and filters steps by z-score, converts each trace into a Reasoning Knowledge Graph with steps as nodes and logical relations as edges, aggregates them into a consensus graph, and synthesizes one new trace by traversing that graph in topological order; across FLD, FOLIO, GSM8K and OlympiadBench with GPT-5.
This work first shows through controlled experiments on PRMBench and ROSCOE that supplying the correct answer does not consistently improve reasoning quality, locating the flaws in the structure of reasoning traces rather than in answer awareness, and then proposes CRAFT: it rolls out K candidate traces per sample, extracts consensus terms via TF-IRF and filters steps by z-score, converts each trace into a Reasoning Knowledge Graph with steps as nodes and logical relations as edges, aggregates them into a consensus graph, and synthesizes one new trace by traversing that graph in topological order; across FLD, FOLIO, GSM8K and OlympiadBench with GPT-5.
This work first shows through controlled experiments on PRMBench and ROSCOE that supplying the correct answer does not consistently improve reasoning quality, locating the flaws in the structure of reasoning traces rather than in answer awareness, and then proposes CRAFT: it rolls out K candidate traces per sample, extracts consensus terms via TF-IRF and filters steps by z-score, converts each trace into a Reasoning Knowledge Graph with steps as nodes and logical relations as edges, aggregates them into a consensus graph, and synthesizes one new trace by traversing that graph in topological order; across FLD, FOLIO, GSM8K and OlympiadBench with GPT-5.