Skip to main content
Back to timeline
arXivSource publication:

Endless Exam scores nine models on 14 families of math construction problems: GPT-6 Astra reaches 143.16 with tools, yet none of 30 published frontiers is surpassed

Synopsis

The authors introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems with 69 evaluation instances, where each submitted object is automatically checked for validity and given an uncapped relative quality score (100 marks average reference parity); across nine models and seventeen configurations, tool-free overall scores span 7.55 to 91.90, and with code and web access GPT-6 Astra, GPT-5.6 Luna and Claude Opus 5.5 reach 143.16, 119.35 and 180.06, while none of the 30 published-frontier references is surpassed.

AI-generated editorial illustration: The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

Interpretation

The benchmark turns construction quality into an automatically verified continuous score: each submitted object is checked against the mathematical constraints and then compared by its objective value (set elements, graph vertices, scalar multiplications) with a published frontier or construction baseline, giving a relative quality where 100 matches the reference on average and improvements beyond it are not clipped. Earlier mathematical reasoning benchmarks mostly score correctness or solve rates, and MathConstruct checks constructed objects while MathConstraint uses solvers to decide feasibility; here a graded quality score on the same instance keeps the magnitude of improvements comparable even after a reference is surpassed. The paper gives the relative quality formula and illustrates it with eight-colour Schur constructions: against the historical reference of 5,041 elements, constructions of 5,286 and 5,362 receive different scores, showing the score distinguishes two advances beyond the reference.

Fourteen parameterised construction families can be scaled by varying parameters, spanning forbidden-pattern sets, geometry, graphs, codes and algebraic designs, with 69 instances in bench-v1.0, of which 30 use published frontiers and 39 use reproducible construction baselines. Unlike a fixed problem set, changing dimension, graph diameter or code length generates harder instances under the same mathematical rules, which the authors describe as one of two routes for continued measurement in the Endless Exam. The paper states the mathematical optimum is unknown for each of the 69 main-evaluation instances and reports problem-size experiments: for maximum-degree-3 graphs the reference rises from 38 at diameter 4 to 1,250 at diameter 10, and for Schur from 160 at five colours to 5,362 at eight colours.

Compact certificates let constructions too large to enumerate still be verified: submissions may use products, lattice orbits, digit sets, difference sets, Cayley generators or algebraic certificates, and the verifier either expands the description or checks the components and algebraic conditions directly. This extends evaluation beyond objects that can be listed element by element; the paper reports that checking the finite factors of a submitted Shannon code certifies over a trillion codewords. Two independently implemented verifiers agree on validity and objective value for all model constructions and all 69 reference constructions, reject all deliberately invalid test cases, and agree with exhaustive checks on small Shannon and trifference instances.

Across nine models, fourteen tool-free configurations and three tool-assisted configurations, the scores separate model size, reasoning effort and tool access, but no model surpasses any of the 30 published-frontier references. Tool-assisted high-effort configurations exceed relative quality 1 on the 39 construction-baseline instances (Astra 1.76, Luna 1.37, Opus 5.5 2.43), showing the score measures improvement beyond a reference while the frontier group remains unbroken. The paper reports tool-free overall scores spanning 7.55 to 91.90; with tools Astra reaches 143.16, Luna 119.35 and Opus 5.5 180.06; on the frontier group Astra, Luna and Opus 5.5 match 29, 23 and 23 references respectively, without surpassing any.

Perspective

The benchmark is aimed at evaluators and model developers who need continued measurement of mathematical construction progress, and applies to construction families whose objective can be computed and whose validity can be checked automatically; the authors invite community contributions of new constructions, task families and verifier improvements, and plan to add larger instances once existing ones become easy while preserving earlier releases for comparability.

The paper states that within the available evaluation budget it tests selected models and reasoning-effort levels, with one evaluation per instance and configuration in the main comparison, so repeated evaluations would be needed to quantify run-to-run variability, which the instance-bootstrap intervals do not measure; answer formats and verification budgets also limit which constructions can be evaluated. In the tool-assisted track, benchmark materials were available online during the Opus 5.5 evaluation but not during the Astra and Luna evaluations, and an audit of Opus 5.5's recorded interactions found no evidence of direct access to them; differing tool interfaces and generation budgets also matter when reading cross-model comparisons.

Sources