Skip to main content

Mathematics

73 items

  1. arXiv

    EAPO couples entropy with the advantage sign for asymmetric credit assignment, topping average accuracy across four backbones

    The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
  2. arXiv

    LSPD recasts on-policy distillation as KL-regularized RL: +1.59 Avg@16 across six math reasoning benchmarks, with a replay variant matching vanilla OPD in about a quarter of rollout batches

    The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
  3. arXiv

    Tsinghua LeapLab replaces KL divergence with a binary directional reward, matching or beating on-policy distillation on math and code, and proposes C-MOPD for multi-teacher training

    This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
  4. arXiv

    Under matched compute on Dream-7B and LLaDA-8B, ORM reranking beats deterministic PRM guidance by up to 82.71% versus 70.02%

    The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
  5. Frontiers in Physics

    SGMR co-evolves strategies and multilayer economic links under one dynamic potential, converging faster with higher welfare and stronger noise resilience on synthetic networks

    This paper develops a co-evolutionary multilayer potential game in which boundedly rational agents update mixed strategies while market, information, and institutional links adapt to observed compatibility and diffusion signals, and proposes a stability-guarded mirror-replicator (SGMR) dynamic combining entropy-regularized strategy revision, projected link rewiring, and a spectral safeguard; the authors prove that the mirror step recovers replicator dynamics in the small-step limit and establish the exact-potential property, monotone potential improvement, sublinear stationarity, and local input-to-state stability under observation disturbances, with computational experiments on synthetic economic networks showing faster convergence, higher collective welfare, stronger noise resilience, an
  6. Frontiers in Applied Mathematics and Statistics

    A game between two hybrid pension managers under jump-diffusion liabilities: competition pushes risk-taking to K=(μ−r)/σ², while own liability jumps lower welfare and the rival's jumps raise it

    The study models two competing hybrid DC pension fund managers as a stochastic differential game in which liability risk follows a jump-diffusion process with truncated exponential jump amplitudes and the exact asset-to-liability ratio formulation is used, and it derives closed-form Nash equilibrium portfolio strategies and value functions via Hamilton–Jacobi–Bellman dynamic programming, finding that equilibrium weights are independent of liability jump parameters, reduce to K=(μ−r)/σ² and are independent of both managers' risk aversion in the symmetric liability-correlation case, that competition strictly amplifies risk-taking relative to the single-agent benchmark, and that a manager's own liability jumps reduce welfare while the competitor's jumps improve it.
  7. CAUCHY Jurnal Matematika Murni dan Aplikasi

    Linear model of coregionalization and co-kriging across 38 Indonesian provinces: moderate spatial dependence (SDP = 68.78%) with the highest predicted stunting in the east

    Using 2024 data for 38 Indonesian provinces on stunting prevalence and nine determinants (low birth weight, safe drinking water, Human Development Index, mean years of schooling, poverty rate, exclusive breastfeeding coverage, mean maternal age at first marriage, antenatal care coverage, and pediatric health service coverage), the study jointly fitted direct and cross semivariograms under a linear model of coregionalization, selected LMC = Nug + Sph(600km), and found moderate spatial dependence of stunting (SDP = 68.
  8. arXiv

    SAGE combines algebraic sparsification and hyperbolic structural guidance to curb long-horizon reasoning biases, beating baselines across 12 benchmarks and 7 model families with up to 8-fold gains on Andrews-Curtis

    The work introduces Symbolic Closure Analysis (SCA) to characterize exploration bias and compounding bias in long-horizon reasoning, and builds SAGE, a framework that injects structural priors during post-training via algebraic sparsification and hyperbolic structural guidance, outperforming SFT, GRPO, EMPO, and GRPO-PRM across 12 benchmarks and 7 model families and achieving up to an 8-fold improvement in Lean-verified proofs on the open Andrews-Curtis task.
  9. arXiv

    LiveMathematicianBench builds a research-level math reasoning benchmark from fresh arXiv papers, where the best model reaches only 43.5% and Gemini-3.1-pro-preview falls to 17.6% under substitution-resistant evaluation

    The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs, introducing a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism; evaluation shows the benchmark is far from saturated, with the best model Gemini-3.1-pro-preview reaching only 43.5%, substitution-resistant evaluation yielding a top score of 30.6% for GPT-5.4 while Gemini-3.1-pro-preview drops to 17.6% below the 20% random baseline, and a dual-mode protocol showing consistent accuracy gains from proof-sketch access.
  10. arXiv

    DMA approximates factor-to-variable messages directly, bounds marginal KL by message KL, and yields a Bayesian neural network inference algorithm without a learning-rate hyperparameter

    The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod

Page 3 · showing 10