Public articles linked to the same research event.
arXiv FERPO is an on-policy maximum-entropy reinforcement learning algorithm that derives an optimal target action distribution from a policy-improvement objective regularized by entropy and KL divergence, then fits the actor to that target by minimizing a forward-KL objective estimated with self-normalized importance sampling (SNIS), thereby performing policy improvement using critic values without differentiating the critic with respect to actions; experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains, and computational benchmarks show faster actor updates than REPPO.
FERPO is an on-policy maximum-entropy reinforcement learning algorithm that derives an optimal target action distribution from a policy-improvement objective regularized by entropy and KL divergence, then fits the actor to that target by minimizing a forward-KL objective estimated with self-normalized importance sampling (SNIS), thereby performing policy improvement using critic values without differentiating the critic with respect to actions; experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains, and computational benchmarks show faster actor updates than REPPO.
FERPO is an on-policy maximum-entropy reinforcement learning algorithm that derives an optimal target action distribution from a policy-improvement objective regularized by entropy and KL divergence, then fits the actor to that target by minimizing a forward-KL objective estimated with self-normalized importance sampling (SNIS), thereby performing policy improvement using critic values without differentiating the critic with respect to actions; experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains, and computational benchmarks show faster actor updates than REPPO.
FERPO is an on-policy maximum-entropy reinforcement learning algorithm that derives an optimal target action distribution from a policy-improvement objective regularized by entropy and KL divergence, then fits the actor to that target by minimizing a forward-KL objective estimated with self-normalized importance sampling (SNIS), thereby performing policy improvement using critic values without differentiating the critic with respect to actions; experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains, and computational benchmarks show faster actor updates than REPPO.