Skip to main content
Back to timeline
arXivSource publication:

Reward Inflation: Gradually Scaling Rewards During Training Speeds Up RL and Suppresses Dormant Neurons

Related research and updates

Synopsis

The work proposes reward inflation, a gradual scaling of rewards over the course of training, and shows theoretically that it induces implicit recency weighting that upweights recent transitions in policy updates while sustaining gradient signals to suppress dormant neurons and help preserve plasticity; experiments on ALE games and MuJoCo tasks indicate that an appropriate level of reward inflation benefits a broad range of tasks, and the authors introduce Fed, an adaptive variant that adjusts the inflation level on the fly and often improves upon fixed inflation.

Source-provided article image: Reward Inflation: A Healthy Stimulus for Reinforcement Learning
Figure 1 ·

Figure 1: (a) The nominal and real performance growth over training time under 2 % 2\% inflation. (b) The inflation rate and inflation level I ⁡ ( g ) I(g) over the training time. (c) The TD error w.r.t. the transition age in replay buffer, across the inflation rates. (d) The changes of gradient norms during training.

arXiv

Interpretation

It proposes reward inflation, a mechanism that gradually scales rewards over the course of training rather than holding reward magnitudes fixed throughout training as is typical. Prior work largely treats reward magnitude as fixed, leaving temporal modulation of rewards underexplored; this work introduces gradual reward scaling during training as a distinct, tunable dimension of the learning signal. At the abstract level it provides a mechanism definition and theoretical argument, and reports empirical results on ALE games and MuJoCo tasks showing that an appropriate level of inflation benefits a broad range of tasks; specific experimental scale and numbers are not given in the abstract.

The theoretical analysis indicates that reward inflation induces implicit recency weighting, so policy updates place more weight on recent transitions, enabling faster adaptation. This offers a mechanistic explanation for scaling rewards: its effect is not merely an overall change in signal strength but a shift in the relative weight of transitions at different times within policy updates. This is a theoretical argument; the abstract states the conclusion but does not provide derivation details or assumptions.

By sustaining gradient signals and mitigating policy saturation, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. It links temporal modulation of rewards to two training-dynamics phenomena, plasticity preservation and dormant neurons, suggesting a pathway by which reward scaling affects internal network state. The abstract gives a theoretical account and reports corroborating empirical results on ALE and MuJoCo; it does not report quantitative measures such as dormant-neuron fractions.

It introduces Fed, an adaptive variant that adjusts the inflation level on the fly, and finds that it often improves upon fixed inflation. It extends fixed inflation to adaptive adjustment of the inflation level in response to training dynamics, reducing the need to hand-set an inflation level. The abstract reports that Fed often improves upon fixed inflation, an empirical comparison; the abstract does not give the specific tasks, baselines, or statistical details of that comparison.

Perspective

The work targets the reinforcement learning training process and applies to settings where reward is the primary learning signal, with empirical support on ALE games and MuJoCo tasks; its conclusions concern an appropriate level of reward inflation, meaning the inflation level must fall within a suitable range. For researchers and practitioners seeking faster policy adaptation, relief from policy saturation and dormant neurons, and preserved network plasticity, this mechanism offers a way to modulate reward magnitude during training; Fed further targets settings where one prefers not to hand-set a fixed inflation level, attempting online adjustment.

The abstract does not give the assumptions and scope of the theoretical derivation, nor the number of tasks, baselines, inflation-level values, or statistical significance in the experiments, so the precise definition of an appropriate level, the magnitude of Fed's improvement over fixed inflation, and the robustness of the conclusions across algorithms and task distributions remain open questions to confirm in the full text. In addition, this summary is based only on the abstract, without the body, figures, or appendix, so implementation details and ablation results cannot be assessed.

Sources