Public articles linked to the same research event.
arXiv The work proposes reward inflation, a gradual scaling of rewards over the course of training, and shows theoretically that it induces implicit recency weighting that upweights recent transitions in policy updates while sustaining gradient signals to suppress dormant neurons and help preserve plasticity; experiments on ALE games and MuJoCo tasks indicate that an appropriate level of reward inflation benefits a broad range of tasks, and the authors introduce Fed, an adaptive variant that adjusts the inflation level on the fly and often improves upon fixed inflation.
The work proposes reward inflation, a gradual scaling of rewards over the course of training, and shows theoretically that it induces implicit recency weighting that upweights recent transitions in policy updates while sustaining gradient signals to suppress dormant neurons and help preserve plasticity; experiments on ALE games and MuJoCo tasks indicate that an appropriate level of reward inflation benefits a broad range of tasks, and the authors introduce Fed, an adaptive variant that adjusts the inflation level on the fly and often improves upon fixed inflation.
The work proposes reward inflation, a gradual scaling of rewards over the course of training, and shows theoretically that it induces implicit recency weighting that upweights recent transitions in policy updates while sustaining gradient signals to suppress dormant neurons and help preserve plasticity; experiments on ALE games and MuJoCo tasks indicate that an appropriate level of reward inflation benefits a broad range of tasks, and the authors introduce Fed, an adaptive variant that adjusts the inflation level on the fly and often improves upon fixed inflation.
The work proposes reward inflation, a gradual scaling of rewards over the course of training, and shows theoretically that it induces implicit recency weighting that upweights recent transitions in policy updates while sustaining gradient signals to suppress dormant neurons and help preserve plasticity; experiments on ALE games and MuJoCo tasks indicate that an appropriate level of reward inflation benefits a broad range of tasks, and the authors introduce Fed, an adaptive variant that adjusts the inflation level on the fly and often improves upon fixed inflation.