EAPO couples entropy with the advantage sign for asymmetric credit assignment, topping average accuracy across four backbones
Synopsis
The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
Interpretation
The authors first run a resampling analysis: on 20 moderately difficult math problems each for Qwen3-4B/8B-Base, they randomly pick two correct and two incorrect responses per problem, fix the prefix before matched-relative-position high- and low-entropy 32-token windows, and sample 64 continuations, finding that originally correct responses are much less likely to succeed again from high- than low-entropy windows (e.g., Qwen3-4B-Base reward 0.390 vs 0.578), while incorrect responses tend to stay incorrect from low-entropy windows (0.070) but reach higher accuracy from high-entropy windows (0.163). Prior work treated high-entropy positions broadly as branching points for exploration but did not separate whether success under uncertainty is reproducible or whether confident failure can be avoided; this analysis characterizes the two cases separately via continuation overlap and final-answer accuracy. Evidence comes from controlled resampling on two models, 20 problems each, two correct and two incorrect responses per problem, and 64 continuations per window, with relative window positions matched to control where resampling starts; the authors also report that high- and low-entropy windows start near the middle of responses with no systematic ordering.
The authors further compare normalized word frequencies between the top and bottom 10% entropy tails of each correct and incorrect response across 2,376 Qwen3-4B-Base math responses, finding exploratory expressions such as "let's" and "consider" enriched at high-entropy positions of correct responses, conclusion markers such as "finally" and "confirm" more prevalent at low-entropy positions of incorrect responses, and exploratory expressions such as "try" and "instead" also appearing at high-entropy positions of incorrect responses. This grounds the resampling-level statistics in observable language markers, indicating that failed responses still retain alternative paths for recovery rather than differing only in overall accuracy. Evidence is a word-frequency comparison over 2,376 responses from a single model, a correlational description that the authors use to motivate rather than prove asymmetric credit assignment.
EAPO normalizes token entropies by batch-level 10th and 90th percentiles with clipping as a detached credit signal, then redistributes the response advantage across tokens via signed uncertainty: favoring high-entropy decisions under positive advantage and low-entropy decisions under negative advantage, with weights normalized by the within-response mean so the advantage's within-response mean and sign are preserved and the ratio between any two weights is bounded by the hyperparameter β. Unlike EntropyAdv's nonnegative entropy bonus, HAPO's sign-preserving entropy adjustment, or 80/20's update restricted to the highest-entropy 20% of tokens, EAPO reverses the entropy preference between reinforcement and penalization and uses only the entropy and advantage already present in rollouts, requiring no auxiliary models, token-level supervision, or privileged information. The method is a deterministic reweighting; the authors provide two propositions with proofs on KL-regularized uncertainty allocation (Proposition 1) and concentration of unchosen alternatives under negative updates (Proposition 2), plus a binary-decision proposition (Proposition 3) in which feedback information from success and failure moves in opposite directions with entropy.
On six competition-level math benchmarks (AIME24/25/26, HMMT26, AMC23, MATH500-H, 297 problems total), EAPO attains the highest mean avg@32 and pass@32 across four backbones: 31.0% and 34.0% mean accuracy with Qwen3-4B-Base and Qwen3-8B-Base, exceeding the strongest entropy-based baselines by 5.6 and 4.3 percentage points; 72.4% and 74.3% with Qwen3-4B and Olmo-3-7B-Think-DPO; and macro averages of 34.89% and 43.02% on eight logical, spatial, and algorithmic Reasoning Gym tasks, while holding the highest pass@k on AIME26 across sampling budgets from 1 to 256. Gains appear on both base and reasoning backbones, on in-distribution math and out-of-distribution reasoning tasks, and ablations show the combination of high-entropy reinforcement with low-entropy penalization is best (6.52 percentage points over uniform credit on Qwen3-4B-Base), indicating the benefit comes from the asymmetric direction rather than from raising entropy or lengthening responses alone. Main results compare avg@32/pass@32 across four backbones, six benchmarks, and 32 samples per problem against five RLVR baselines and the initial backbone; robustness evidence includes matched-token-budget continuation experiments, scaling from 1.7B to 14B, LoRA versus full fine-tuning, and training-reward curves over three random seeds.
Perspective
The result targets research and engineering settings that train LLM reasoning with verifiable rewards: EAPO needs only the policy's own entropy and the existing response advantage, so it drops into GRPO-style pipelines, applies to both base and reasoning backbones, and is validated on competition math problems and on logical, spatial, and algorithmic Reasoning Gym tasks. For a reader, this means an off-the-shelf option for improving exploration without auxiliary models, extra sampling, or privileged information; the authors also note that the allocation depends on how well the model's own uncertainty aligns with the true contribution of individual decisions, so it fits settings where model uncertainty still reflects reasoning branches. The authors list combining findings across states or trajectories as future work and point to scientific-discovery practices that maintain diverse candidates and iteratively reuse them.
Entropy as a credit signal reflects the model's own learned preferences and confidence, and how well it aligns with the true contribution of individual reasoning decisions remains an open question; the mechanism analysis focuses on alternative continuations from a given reasoning state, leaving combination of findings across states or trajectories unexplored. Ablations show larger β generally improves faster but is less stable, and the default is a trade-off between stability and concentration. Main results are reported as avg@32 and pass@32, training is primarily LoRA on DAPO-Math-17k-Processed, and full fine-tuning and multiple-seed results serve as supplements; behavior beyond these settings is for readers to judge against their own tasks.
