LiveMathematicianBench builds a research-level math reasoning benchmark from fresh arXiv papers, where the best model reaches only 43.5% and Gemini-3.1-pro-preview falls to 17.6% under substitution-resistant evaluation
Synopsis
The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs, introducing a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism; evaluation shows the benchmark is far from saturated, with the best model Gemini-3.1-pro-preview reaching only 43.5%, substitution-resistant evaluation yielding a top score of 30.6% for GPT-5.4 while Gemini-3.1-pro-preview drops to 17.6% below the 20% random baseline, and a dual-mode protocol showing consistent accuracy gains from proof-sketch access.
Figure 1: Overall accuracy on LiveMathematicianBench across different LLMs. Even recent frontier models remain far from saturation. Random: 20%.
arXivInterpretation
It introduces LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs. Unlike existing benchmarks limited by synthetic settings and data contamination, it grounds evaluation in newly published theorems, providing a realistic testbed beyond memorized patterns. The abstract states the benchmark is built from arXiv papers published after training cutoffs and describes it as a scalable, contamination-resistant testbed; the number of items and construction details are not given in the loaded text.
It introduces a thirteen-category logical taxonomy of theorem types (e.g., implication, equivalence, existence, uniqueness) for fine-grained evaluation across reasoning forms. Existing benchmarks typically do not stratify by the logical type of theorems, so this taxonomy enables distinguishing performance across different reasoning forms. The abstract explicitly gives the count of thirteen categories and examples including implication, equivalence, existence, and uniqueness.
It employs a proof-sketch-guided distractor pipeline that uses high-level proof strategies to construct plausible but invalid answer choices reflecting misleading proof directions. Compared with surface-matching option design, this pipeline increases sensitivity to genuine understanding over surface-level matching. The abstract describes the construction principle but does not report the number of distractors or human validation details.
It introduces a substitution-resistant mechanism to distinguish answer recognition from substantive reasoning, and evaluation shows sharp accuracy drops plus gains from proof sketches. The mechanism separates recognition-style answering from genuine reasoning, and the dual-mode protocol further shows models can leverage high-level proof strategies. The abstract reports the best model Gemini-3.1-pro-preview at 43.5%; under substitution-resistant evaluation GPT-5.4 scores highest at 30.6% while Gemini-3.1-pro-preview falls to 17.6%, below the 20% random baseline; the dual-mode protocol shows consistent accuracy gains from proof-sketch access.
Perspective
The benchmark targets research-level mathematical reasoning and applies to dynamic evaluation settings where items are drawn from newly published theorems; its audience is researchers and evaluation designers studying LLM mathematical and reasoning capabilities. Its contamination-resistant and scalable design makes it suitable for reuse across model iterations, and the proof-sketch mode allows examining how models exploit high-level proof strategies.
The loaded text is an abstract and does not provide the number of items, the criteria for selecting source papers, the concrete implementation of the distractor and substitution-resistant mechanisms, or confidence intervals and run-to-run variance; how stable these evaluation numbers are across item subsets and runs, and how performance distributes across the thirteen logical categories, remain open questions that require the full paper.
