Bellman Policy Optimization: Rewriting Policy Mirror Descent as a Critic-Free Trajectory-Level Objective via the Bellman Equations
Synopsis
The work introduces Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD): for autoregressive generation with terminal rewards, the Bellman equations reformulate PMD into a trajectory-level objective that avoids estimating intermediate state values, and it is proved that this objective shares the same unique optimal solution as the original PMD objective; the practical loss replaces GRPO's importance-sampling ratio with a smoothed ratio of complementary token probabilities as a mismatch-correction weight, achieving higher average accuracy than GRPO-ClipHigher, GSPO, CISPO, and DPPO on mathematical reasoning benchmarks.
Interpretation
It proposes and proves a critic-free, trajectory-level reformulation of PMD: under terminal rewards, the Bellman equations express each token advantage as a difference between value functions at consecutive states, and these differences telescope along a response, leaving only the terminal reward and the initial value, so the objective depends only on the terminal reward and the initial value function rather than on value or advantage estimates at intermediate states. Whereas directly implementing PMD requires estimating values at every visited intermediate state (typically via a separately trained value model that increases memory and compute cost, with learned value estimates that can be inaccurate on reasoning tasks), this reformulation removes value estimation from intermediate states; relative to critic-free methods such as GRPO and its variants, it supplies a derivation path starting from PMD. The paper states Theorem 1 with sufficiency and necessity proofs: on states reachable under the rollout policy, the reformulated objective and the original PMD objective have the same unique optimal solution, and the equivalence holds for any positive weighting function; the necessity proof uses an argument in which token-level residuals form a martingale under the rollout policy.
Starting from this trajectory-level objective, the practical BPO loss is obtained by linearizing the squared-residual objective, estimating the initial value and a prompt-dependent scaling factor from grouped rollouts, approximating the full reverse KL divergence with binary KL divergence, and applying additive smoothing, masking, and clipping; its mismatch-correction weight is a smoothed ratio of complementary token probabilities, replacing GRPO's token-level importance-sampling ratio. GRPO uses a PPO-style clipped surrogate objective with token-level importance-sampling ratios; BPO retains the GRPO per-token loss form but replaces the importance-sampling ratio with a truncated mismatch-correction weight determined by the rollout and current token probabilities and additively smoothed for numerical stability. The derivation proceeds in four steps (linearized approximation, group-based estimation and normalization, binary KL approximation with additive smoothing, masking and clipping), with the binary KL identity proved in Appendix A.2; the paper also notes that binary KL divergence provides a lower bound for KL divergence.
Trained on Qwen3-30B-A3B-Base with the English subset of DAPO-Math-17k and evaluated on AIME24, AIME25, and AIME26 using Avg@32 as an estimate of Pass@1, BPO reaches 50.5% average accuracy, above GRPO-ClipHigher's 39.5% and the strongest baseline CISPO's 47.4%, corresponding to gains of 11.0 and 3.1 percentage points, and is highest on all three benchmarks. In a controlled comparison that varies only the policy loss while keeping all other experimental settings identical, BPO achieves higher average accuracy than all four baselines GRPO-ClipHigher, GSPO, CISPO, and DPPO; after 400 training steps its average accuracy is 49.4%, versus 45.5% for DPPO, the strongest baseline at the end of training. Each rollout batch contains 256 prompts with 16 responses per prompt, and the resulting 4096 responses are split into 8 minibatches of 512; training runs 400 steps corresponding to 3200 optimizer updates, with a maximum response length of 16384 tokens, and all methods use rollout-router replay (R3); Table 1 reports each method's results at the checkpoint with the highest mean accuracy across the three benchmarks.
Ablations on Qwen3-4B-Base with 1000 training steps show that BPO performs similarly across a range of smoothing and truncation settings and exceeds the GRPO-ClipHigher baseline in all of them. The ablations show BPO is insensitive to the smoothing parameter over a range of values (average accuracy 25.4%-25.8%) and yields average accuracies between 25.3% and 25.8% across three values of the truncation parameter, a spread of 0.5 percentage points, all above GRPO-ClipHigher's 20.5%. The ablations use 128 prompts per batch with 16 responses per prompt (2048 responses split into four minibatches of 512), a maximum response length of 8192 tokens, and each sweep changes only one hyperparameter while sharing the same GRPO-ClipHigher baseline.
Perspective
The result targets autoregressive generation with terminal rewards and reinforcement learning with verifiable rewards; experiments train Qwen3-30B-A3B-Base on the English subset of DAPO-Math-17k and evaluate on mathematical reasoning benchmarks (AIME24, AIME25, AIME26), with ablations on Qwen3-4B-Base, and the theoretical equivalence is restricted to states reachable under the rollout policy. This provides a reusable derivation framework and loss form for researchers and practitioners who want to improve policy optimization without training a value model, provided the task supplies terminal verifiable rewards.
A careful reader would still watch how the equivalence and the loss behave beyond mathematical reasoning, on other base models, and at longer training scales; how the smoothing and truncation parameters of the mismatch-correction weight behave on larger models; and that some numbers appear in placeholder form in the abstract and body, so exact figures should be checked against the original tables. In addition, this reading is full text, but specific curve values in the figures should be taken from Figure 1 of the original.
