Mathematics
72 items
RIDE extrapolates RL teachers' hidden-state residuals and matches or beats the teacher on average across four base/teacher pairs
The work proposes RIDE (RL-Induced Direction Extrapolation): on student-generated trajectories it computes the layerwise hidden-state residual between an RL-trained teacher and its pre-RL checkpoint, then regresses the student's hidden states toward targets displaced beyond the teacher along that residual; across four base/RL-teacher pairs (R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, Phi-4-mini) RIDE's mean Avg@16 on AIME24, AIME25 and AIMO approaches or exceeds the RL-trained teacher on every pair, the only compared method whose mean does so, and it consistently outperforms output-space extrapolation.
EvoDuet lets LLMs retrieve the web by knowledge gap during evolutionary search, raising OpenEvolve's normalized discovery gain from 74.1% to 78.0% (GPT-5.6-Luna) and from 61.3% to 82.3% (Gemini-3.8-Flash) across 21 optimization tasks and surpassing prior best scores on eight
The work introduces EvoDuet, a bi-level method that co-evolves solutions and web search queries with fixed model parameters: at each iteration a knowledge-gap-based retrieval gate decides whether to retrieve new documents, reuse stored ones, or proceed without them, an inner loop refines queries and ranks documents by the solution scores they are predicted to yield, and an outer loop generates candidates in parallel and records evaluated outcomes for later searches; across 21 optimization tasks with one candidate per iteration it raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, Qwen3.
Endless Exam scores nine models on 14 families of math construction problems: GPT-6 Astra reaches 143.16 with tools, yet none of 30 published frontiers is surpassed
The authors introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems with 69 evaluation instances, where each submitted object is automatically checked for validity and given an uncapped relative quality score (100 marks average reference parity); across nine models and seventeen configurations, tool-free overall scores span 7.55 to 91.90, and with code and web access GPT-6 Astra, GPT-5.6 Luna and Claude Opus 5.5 reach 143.16, 119.35 and 180.06, while none of the 30 published-frontier references is surpassed.
OASIS restricts self-distillation supervision to verified trajectories, widening its average gain over OPSD from 0.59 to 3.05 points across Qwen3-1.7B, 4B and 8B
Using truncated-continuation probes, the work finds that in on-policy self-distillation (OPSD) the teacher's advantage over the student is concentrated on failed trajectories (77% versus 37% recovery at 1.7B) and shrinks with scale, and therefore proposes OASIS: keeping the OPSD objective unchanged, it applies supervision only along the shortest outcome-verified on-policy trajectory and replaces the written reference solution with a distinct same-problem self-generated rollout, typically the student's own unverified attempt, improving mean Avg@12 over OPSD by 0.59, 1.86 and 3.05 points on Qwen3-1.7B, 4B and 8B across AIME 2024, AIME 2025 and HMMT 2025 while requiring no written solutions.
LANTERN used Qwen3-32B hidden states to rank 50 million OEIS sequence pairs and returned 62 verified relations in under 8 hours, four of them absent from OEIS and a targeted literature search
LANTERN trains a linear classifier on Qwen3-32B activations to rank 50 million candidate relations among the 10,000 most-referenced OEIS sequences, then applies staged filtering, hypothesis generation, executable verification and analytical checking to produce 62 verified relations without an existing OEIS cross-reference; a content screen retained 13, nine of them informative or insightful, four not found in the OEIS or a targeted literature search, with the end-to-end process taking under 8 hours.
LSD cuts the length-scaling tax on already-solved queries from 19.0% to -3.7% while matching or beating RL Pass@1
The work defines the length-scaling tax (LST) as the excess response length that RL post-training adds on already-solved queries without a matching accuracy gain, and proposes Length Self-Distillation (LSD), which routes solved prompt groups to on-policy distillation while keeping the original RLVR objective for unsolved groups, using an exponential moving average of the same policy lineage as teacher with no external model; LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks while matching or improving average Pass@1 over RL.
S²D-OPD keeps only the top 10% of states per response by teacher–reference JSD, lifting held-out accuracy from 46.36% to 47.31% in seven of eight teacher–student settings
The work shows that Direct-OPD's token-level log-ratio reward measures only relative change and can stay fixed while the probability mass that both checkpoints assign to the student's candidate tokens vanishes, whereas the teacher–reference JSD and both KL directions vanish with that mass; it therefore proposes S²D-OPD, which ranks student-sampled states by teacher–reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% per response, improving held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings across two teacher pairs and four student models from 1.7B to 8B parameters and matching it in the eighth, without extra forward passes.
Rachel Webb argues LLMs will drastically change how she executes math research but not her metric for mathematical interest or her two humanistic reasons for doing math.
In this guest post, Rachel Webb draws on her own mathematical research experience to argue that LLMs let her execute research faster and turn some of her lands of mathematical fantasy into worlds she can realistically start exploring, while her metric for mathematical interest stays unchanged and humans keep doing math for two humanistic reasons: math is interesting to us individually, and math creates communities.
Delaunay-weighted two-sample test uses geometric direction information to detect principal-direction covariance differences in high-dimensional manifold data
The authors propose a Delaunay-weighted two-sample test: under a low-dimensional manifold assumption they define a Delaunay weight from the Delaunay triangulation that captures both geodesic distance and relative direction, use the average within-group Delaunay weight as the test statistic with a permutation p-value, prove asymptotic normality under the null and consistency under the alternative, and show in simulations substantially higher power than k-NN, k-MST, kernel, e-distance, covariance, and regression tests when the two distributions differ in the principal directions of their covariance matrices, while detecting a treatment-group difference with p=0.011 in a mice protein expression dataset.
Page 1 · showing 10