Skip to main content

Mathematics

73 items

  1. arXiv

    Bellman Policy Optimization: Rewriting Policy Mirror Descent as a Critic-Free Trajectory-Level Objective via the Bellman Equations

    The work introduces Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD): for autoregressive generation with terminal rewards, the Bellman equations reformulate PMD into a trajectory-level objective that avoids estimating intermediate state values, and it is proved that this objective shares the same unique optimal solution as the original PMD objective; the practical loss replaces GRPO's importance-sampling ratio with a smoothed ratio of complementary token probabilities as a mismatch-correction weight, achieving higher average accuracy than GRPO-ClipHigher, GSPO, CISPO, and DPPO on mathematical reasoning benchmarks.
  2. Terence Tao blog RSS

    Open problems, open mathematics

    This guest post by Antonio Auffinger draws on his decade of conversations with biologists, computer scientists, and physicists to argue that mathematics' culture of open sharing may erode if machine proof generation becomes fast and accessible while credit systems remain unchanged, and it calls on the mathematical community to rethink incentives while embracing biology and applied sciences as sources of new mathematical questions and phenomena.
  3. arXiv

    1% of Tokens Can Match Full Distillation: IER Re-ranks Sparse Supervision by Gradient-Estimation Reliability

    The work reframes token selection in on-policy distillation (OPD) as a gradient-estimation reliability problem at a fixed prefix, decomposes the one-sample reverse-KL gradient into signal and sampling noise in information geometry, proposes the information-efficiency ratio (IER) as a signal-to-noise ratio under the optimal scalar baseline, and approximates it on a top-K candidate set for ranking; on mathematical and medical reasoning, IER alone approaches or exceeds full OPD at token budgets of 0.1%–1%, and combined with existing usefulness scores (IER-OR, IER-AND) it matches or exceeds full OPD in multiple settings.
  4. arXiv

    Bringing survey small-area estimation into AI evaluation: PP-S and PP-TS improve point and interval estimates on a benchmark and on deployed traffic, while DB-CV picks as well as an independent validation sample from one sample

    The work frames disaggregated AI evaluation as finite-population survey sampling, proposes prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's prediction-powered estimate, plus a taxonomy extension (PP-TS) that borrows strength along a nested reporting hierarchy, and derives an approximately unbiased design-based cross-validation score (DB-CV); on the Open LLM Leaderboard (9,324 questions, 34 task types) and PRISM deployed traffic (68,371 rated responses, 21 LLMs by 3 conversation types), where every outcome is observed, the smoothed estimators beat direct estimators in point and interval estimation with near-nominal 95% coverage, and DB-CV selects as well as an independent validation sample at the same budget while estimating the chosen estimator's error far more ac
  5. arXiv

    Making KDA's gate signed lets a single CKDA layer track finite groups and extrapolate periodic waveforms, with 1.3B downstream accuracy on par with KDA

    The work introduces Complex KDA (CKDA): extending Kimi Delta Attention by allowing signed gate entries and an extended delta-rule coefficient range so that a single diagonal-plus-rank-one transition can realize a 2D rotation; the authors prove every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition, show one CKDA layer tracks every finite group isomorphic to a subgroup of O(2), save one layer relative to comparable diagonal-plus-rank-one linear RNNs, and report experiments on group-word problems, periodic audio continuation, and language modeling.
  6. arXiv

    Separating Teacher Self-Deviation from the Distillation Signal: Calibrated On-Policy Distillation

    The work argues that the token-level teacher–student discrepancy used by on-policy distillation (OPD) mixes in context-induced teacher-side variation, termed Teacher Self-Deviation (TSD), and introduces Cal-OPD, which estimates the teacher's self-deviation region via positive and negative privileged interventions and keeps only the residual beyond that region as the optimization signal, consistently outperforming standard OPD and its variants on mathematical reasoning benchmarks while retaining only about 52–65% of the original discrepancy.
  7. Terence Tao blog RSS

    Why Do We Need Human Mathematicians Anymore?

    This guest opinion piece by Po-Shen Loh, published on Terence Tao's blog, proposes that the mathematics community and all industries should publicly adopt the axiom 'We (humans) should help humanity flourish,' and argues from it that further AI advance will create more high-skill human oversight jobs than there are people to fill them, eventually forcing AI progress to slow; it also discusses how pure mathematics contributes to human flourishing and what practical changes adopting the axiom might bring.
  8. Terence Tao blog RSS

    If math is more than proof, we need to better celebrate the rest of it

    This is a guest opinion piece by Grant Sanderson published on Terence Tao's blog, arguing that the mathematics community should more firmly define and grant academic credit to a kind of work it calls a "motivated explanation" — exposition that places definitions in the middle, may begin from a relatable but not-quite-right idea, and aims to answer "how would you think of that?" — and offering concrete institutional suggestions such as making the deliverable of a small problem a talk, enumerating unsolved exposition problems, founding journals focused on understanding, and valuing great textbook writing more in hiring and tenure.
  9. Google Research

    MilleMiglia: A realistic instance generator for middle-mile logistics

    The work introduces and open-sources MilleMiglia, a C++ instance generator serialized with Protocol Buffers that synthesizes middle-mile logistics networks from statistical distributions over spatial placement, demand, and vehicle rotations, embedding hard constraints such as fixed schedules, distribution-center throughput limits, and cross-vehicle synchronization into a single data format, thereby offering reproducible benchmarks from small academic toy problems to continent-wide industrial scale while preserving corporate privacy.
  10. Terence Tao blog RSS

    SAIR's Open Math Model Initiative

    In this blog post, Terence Tao announces that the SAIR Foundation is launching an Open Math Model initiative, inviting the mathematical community and its supporters to build open-source models together and openly seeking partners who can contribute funding, compute, expertise, or community building; the initiative sets out principles of community-shaped models, tools for everyday mathematical work (understanding difficult arguments, checking references, exploring examples, writing code, and formalizing proofs), open development (open licensed weights and code, published training methods, reproducible evaluations), community ownership of data, shared intellectual property (for example Apache 2.0, MIT, or CC BY 4.

Page 6 · showing 10