LiveMathematicianBench evaluates LLMs on arXiv theorems published after training cutoffs: best model reaches only 43.5%, and Gemini-3.1-pro-preview falls to 17.6% under substitution resistance, below the random baseline
Synopsis
The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs; it introduces a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism, and evaluation shows the benchmark is far from saturated—the best model, Gemini-3.1-pro-preview, achieves only 43.5%, while under substitution-resistant evaluation GPT-5.4 scores highest at 30.6% and Gemini-3.1-pro-preview falls to 17.6%, below the 20% random baseline; a dual-mode protocol shows proof-sketch access yields consistent accuracy gains.
Figure 1: Overall accuracy on LiveMathematicianBench across different LLMs. Even recent frontier models remain far from saturation. Random: 20%.
arXivInterpretation
It presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs. Compared with existing benchmarks limited by synthetic settings and data contamination, it grounds evaluation in newly published theorems, providing a more realistic testbed beyond memorized patterns. The abstract states the benchmark is built from arXiv papers published after training cutoffs and reports that the best model, Gemini-3.1-pro-preview, achieves only 43.5%, indicating the benchmark is far from saturated.
It introduces a thirteen-category logical taxonomy of theorem types (e.g., implication, equivalence, existence, uniqueness), enabling fine-grained evaluation across reasoning forms. Existing benchmarks typically do not stratify by the logical type of theorems, whereas this taxonomy allows evaluation to distinguish performance across reasoning forms. The abstract explicitly lists thirteen theorem-type categories with examples such as implication, equivalence, existence, and uniqueness, and states the taxonomy enables fine-grained evaluation.
It designs a proof-sketch-guided distractor pipeline that uses high-level proof strategies to construct plausible but invalid answer choices reflecting misleading proof directions. Compared with surface-matching option design, this pipeline increases sensitivity to genuine understanding over surface-level matching. The abstract states the pipeline uses high-level proof strategies to construct distractors and notes its aim of increasing sensitivity to genuine understanding.
It introduces a substitution-resistant mechanism to distinguish answer recognition from substantive reasoning, and uses a dual-mode protocol to examine the role of proof sketches. The mechanism and protocol separate recognition-style answering from substantive reasoning and test whether models can leverage high-level proof strategies for reasoning. The abstract reports substitution-resistant results: GPT-5.4 scores highest at 30.6% and Gemini-3.1-pro-preview falls to 17.6%, below the 20% random baseline; the dual-mode protocol shows proof-sketch access yields consistent accuracy gains.
Perspective
The benchmark targets evaluation of research-level mathematical reasoning and is intended for researchers and evaluation designers who want to test LLM reasoning with newly published theorems; its dynamic construction allows continuous expansion with new papers, the thirteen-category logical taxonomy and proof-sketch-guided distractors support fine-grained analysis by reasoning form, and the substitution-resistant mechanism and dual-mode protocol offer a reusable evaluation setup for separating answer recognition from substantive reasoning.
The abstract does not state the number of benchmark items, the paper selection scope and time window, or the distribution across the thirteen theorem types, nor does it give a full score table for all models in the standard mode; the implementation of the substitution-resistant mechanism and the form in which proof sketches are provided in the dual-mode protocol are likewise not detailed in the abstract. In addition, the abstract does not report human expert or mathematician performance on the same benchmark, so it is difficult to judge what 43.5% and 30.6% mean relative to human levels. These are open questions about evaluation detail and external validity worth watching when reading the full text.
