Skip to main content

Mathematics

73 items

  1. Statistica Sinica

    Delaunay-weighted two-sample test uses geometric direction information to detect principal-direction covariance differences in high-dimensional manifold data

    The authors propose a Delaunay-weighted two-sample test: under a low-dimensional manifold assumption they define a Delaunay weight from the Delaunay triangulation that captures both geodesic distance and relative direction, use the average within-group Delaunay weight as the test statistic with a permutation p-value, prove asymptotic normality under the null and consistency under the alternative, and show in simulations substantially higher power than k-NN, k-MST, kernel, e-distance, covariance, and regression tests when the two distributions differ in the principal directions of their covariance matrices, while detecting a treatment-group difference with p=0.011 in a mice protein expression dataset.
  2. arXiv

    Tsinghua team turns the SIS proposal into a learning problem with GFlowNets: one network zero-shot matches or beats the post-hoc best of 31 analytic proposals on 1,190 unseen margins

    The work shows that the zero-variance sequential importance sampling (SIS) proposal for binary matrices with fixed margins is exactly the policy of a unit-reward GFlowNet, and proposes MarginFlow, a set transformer that reads the remaining margins and, trained on 1,904 margins, runs zero-shot on 1,190 held-out margins, matching or beating the post-hoc best of 31 analytically designed configurations on 1,187 of them with a median effective sample fraction of 99.8%.
  3. arXiv

    SAKI routes teacher supervision through maximal-coupling accept/correct events, lifting Mean@8 and Pass@8 for both 1.7B and 0.6B students across seven math reasoning benchmarks

    SAKI realizes a KL-constrained teacher-guided rollout through maximal coupling and reuses the realized accept/correct events as a token-level supervision router: accepted positions keep sampled-token reverse-KL, correction positions switch to direct supervision on the teacher's highest-probability token, and the correction probability is exactly TV(p_t,q_t) so the same trust-region radius upper-bounds intervention frequency; an engine-resident speculative verifier preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x, and across seven mathematical reasoning benchmarks SAKI improves Mean@8 and Pass@8 over the matched teacher-guided baseline for both 1.7B and 0.6B students.
  4. arXiv

    GRAFT swaps trajectories between two heterogeneous models, beating GRPO for both at equal budget with a 2.1-point average gain

    The work proposes GRAFT, a framework that in Reinforcement Learning with Verifiable Rewards (RLVR) replaces a receiver's all-fail rollout groups with peer groups containing both successful and unsuccessful responses, controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping; across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT improves both models over GRPO at the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points, with 1.8 points on average preserved when reusing stored peer trajectories.
  5. arXiv

    ZJU team runs 25 same-family on-policy distillation pairs on Qwen2.5 from 0.5B to 14B, finding peak capability is predictable from student scale and teacher score, and that weaker teachers can teach stronger students

    The study systematically characterizes the scaling properties of same-family on-policy distillation (OPD) on Qwen2.5 (0.5B–14B) math reasoning, finding a regular early useful-transfer regime in which held-out accuracy rises approximately linearly with the square root of token-level reverse KL from the student initialization, and fitting power laws in student scale, teacher scale, and teacher gold score that predict peak accuracy and transfer rate, with every weak-to-strong student peaking above its own teacher.
  6. Journal of Machine Learning

    OptimAI turns natural-language optimization problems into solver code with a multi-agent LLM pipeline, reaching 88.1% on NLP4LP and 82.3% on Optibench and cutting error rates by 58% and 52% over the prior best

    The work introduces OptimAI, an LLM-powered multi-agent framework that takes a natural-language optimization problem through four stages—formulation, planning, solver code generation, and reflective debugging—and adds UCB-based debug scheduling to switch dynamically among candidate plans; under zero-shot prompting it reaches 88.1% accuracy on NLP4LP with GPT-4o+o1-mini and 82.3% on Optibench with DeepSeek-R1, reducing error rates by 58% and 52% over the prior best, while ablations show that removing the planner or code critic drops productivity by 5.8× and 3.1× and that enabling UCB debug scheduling adds a further 3.3× productivity gain.
  7. CSIAM Transactions on Applied Mathematics

    Chen, Ji and Xu propose the DiGCA phase classifier, generating Lifshitz-Petrich phase diagrams about two orders of magnitude faster with over 98% classification accuracy

    The authors propose a Derivative-informed Graph Convolutional Autoencoder (DiGCA) phase classifier that feeds both the Lifshitz-Petrich model solutions and their derivatives (the nonlocal term G(φ)) into a graph convolutional autoencoder for dimensionality reduction, then classifies with a fully connected neural network, generating phase diagrams over the parameter domain [−0.01,0.05]×[0,1] with over 98% classification accuracy, roughly two orders of magnitude faster than MCMS-RBM, and remaining stable under up to 10% additive white noise.
  8. arXiv

    DCSD decouples credit direction from magnitude in self-distillation, topping 11 benchmarks and lifting math reasoning by 8.45 points over base models

    The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
  9. arXiv

    Swapping matrix multiplication for an associative-algebra product: a 110M-parameter model gains 6.2–7.8% generation throughput while GSM8K, MBPP and IFEval all drop

    The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
  10. arXiv

    CIS recasts training-inference mismatch as a log-odds displacement and tops the five-benchmark average on all three MoE models

    The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat

Page 2 · showing 10