Skip to main content
Back to timeline
arXivSource publication:

TGRL turns temperature differences into a training signal: +1.6% math average, +196.7 CodeForces rating, with no extra rollout budget

Synopsis

The work proposes Temperature-Grouped Reinforcement Learning (TGRL), which partitions each prompt's rollout group into low- and high-temperature subsets, estimates prompt-level exploration gain from their reward contrast, and allocates that gain as token-level credit using Jensen–Shannon divergence between the temperature-scaled next-token distributions induced by the same logits; across 11 benchmarks, TGRL improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%, reaching equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget.

AI-generated editorial illustration: TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs

Interpretation

TGRL turns sampling temperature from a decoding parameter into an explicit training signal: each prompt's rollout group is split into a low-temperature reference subset and a high-temperature exploration subset, and the mean-reward gap between them serves as an unbiased estimate of prompt-level exploration gain that is injected as an additive term into the mixed-group advantage. Prior RLVR exploration approaches either expand the sample budget at rollout time (test-time scaling) or treat exploration as a global regularizer, auxiliary objective, or decoding policy (temperature scheduling); TGRL instead converts temperature-induced diversity into a reward-grounded, prompt-level gain estimate. The paper derives Lemma 1 showing that, with independent low- and high-temperature rewards, the difference of subgroup means is an unbiased estimate of prompt-level exploration gain; Proposition 1 decomposes the mixed-group advantage into a within-subgroup fluctuation term plus a cross-group shift controlled by the exploration gain, supported by a synthetic diagnostic.

The exploration gain is allocated across tokens by JS divergence: Jensen–Shannon divergence between the low- and high-temperature next-token distributions from the same logits gives larger credit to more temperature-sensitive positions, concentrating gradient updates there. Earlier methods broadcast the exploration signal uniformly over the response or keep it as a sequence-level scalar; TGRL provides a token-level allocation criterion and proves that any monotone JS weighting assigns more mass than uniform weighting to temperature-sensitive positions. Lemma 2 bounds token-level JS by the path-averaged logit variance, and Proposition 2 establishes the monotone JS credit-allocation result; a local-cooling probe shows JS-selected spans yield the highest flip rate (12.1%) and largest mean score drop (0.241), above Low-Margin (11.6%, 0.232), Matched-Random (8.7%, 0.173), and Entropy (7.6%, 0.152).

Across 11 benchmarks, TGRL broadly improves over strong RLVR baselines without expanding the rollout budget, and training dynamics indicate the gain comes from token-level credit allocation rather than reward scale. Gains concentrate on competition-level math, unsaturated code benchmarks, and sparse-reward long-horizon agent tasks, while remaining competitive on near-saturated benchmarks, consistent with a design that judges per prompt whether broader exploration helps. The six-benchmark math average is 70.2 at 32B (strongest baseline 68.6) and 69.4 at 14B (strongest baseline 68.0); CodeForces rating rises from 1224.0 to 1574.3 (+196.7) and LiveCodeBench Pass@16 improves by 4.4%; ALFWorld overall success reaches 86.7% (+6.3% over the best RL baseline) and WebShop success 74.2% (+4.9%); TGRL and the uniform-token-credit variant have nearly identical late-stage training rewards (0.394 vs. 0.388) yet differ in downstream accuracy, response length, and truncation.

Wall-clock experiments show TGRL learns faster under the same compute budget: it attains a higher AIME-Combined score than GRPO within a matched 12-hour budget and reaches DAPO's accuracy target in about 6.39 hours versus DAPO's about 10.01 hours. The paper links exploration-gain estimation to learning progress per unit wall-clock, showing that mixed-temperature rollout construction and token-wise JS post-processing do not translate into a visible wall-clock penalty. Complexity analysis (Proposition 4, Corollary 1) shows TGRL shares the leading asymptotic order of single-temperature group optimization, with only a lower-order token-wise JS term on the high-temperature group; wall-clock comparisons use the same hardware, data pipeline, and decoding configuration.

Perspective

The results target reinforcement learning with verifiable rewards: mathematical reasoning, code generation, and benchmarked long-horizon agent tasks, trained on DeepScaleR, DeepCoder, and the WebShop/ALFWorld environments with Qwen3-4B/14B/32B and Qwen2.5-7B-Instruct backbones. The default configuration uses 4 rollouts per prompt, a 1:3 split between low temperature 0.3 and high temperature 1.2, and a 20-step single-temperature warmup, with the low-temperature group contributing only to group statistics and receiving no token-level gradient. The authors explicitly do not claim the mechanism directly solves open-ended dialogue, subjective preference optimization, safety-critical decision making, or tasks with noisy, delayed, or hard-to-verify rewards. Future work may extend the principle to adaptive multi-temperature grouping, richer token-credit allocation criteria, and more open-ended interactive environments.

Competition-level math benchmarks are small (30 problems each in AIME24/AIME25, 40 in AMC23); the paper reports 5-seed Avg@16 evaluations with cross-seed standard deviations typically 1–4 percentage points and states that mean gaps on AIME exceed per-seed standard deviations. Readers may still watch behavior across different backbones, reward verifiers, and longer training. The temperature-sensitivity study shows performance degrades mainly when the temperature gap is too narrow, with a broad competitive regime at larger gaps, though the best temperature still varies by task and checkpoint. The limitations section notes the mechanism does not cover open-ended dialogue, subjective preferences, safety-critical decisions, or noisy or delayed rewards. In addition, the homepage evidence bundle's external story has some numeric gaps in its rendering (for example, missing digits where 'up to faster' and 'by points' appear); the specific numbers in this summary follow the readable figures in the paper body and tables.

Sources