Skip to main content
Back to timeline
arXivSource publication:

FERPO optimizes maximum-entropy policies without critic action gradients, showing competitive performance and sample-efficiency gains on MuJoCo Playground and ManiSkill

Related research and updates

Synopsis

FERPO is an on-policy maximum-entropy reinforcement learning algorithm that derives an optimal target action distribution from a policy-improvement objective regularized by entropy and KL divergence, then fits the actor to that target by minimizing a forward-KL objective estimated with self-normalized importance sampling (SNIS), thereby performing policy improvement using critic values without differentiating the critic with respect to actions; experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains, and computational benchmarks show faster actor updates than REPPO.

Source-provided article image: FERPO: Forward Entropy-Regularized Policy Optimization
Figure 1 ·

Figure 1: Actor directions and local losses. Top: green/red arrows point toward/away from the cost minimum (star); lengths are illustrative, not convergence guarantees. Bottom: frozen-reference losses at A–C, zeroed at the reference; KL losses are scaled by T T . Details: Appendix A.1 .

arXiv

Interpretation

FERPO performs policy improvement using critic values rather than action gradients of the critic, avoiding the step where accurate value predictions need not imply accurate action derivatives. Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic, whereas FERPO explicitly removes action derivatives from the policy-improvement path. The abstract states this design motivation argumentatively and describes FERPO as an on-policy maximum-entropy algorithm; implementation details and ablation settings require the full text.

FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and KL divergence, and fits the actor with a forward-KL objective. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. The abstract gives the derivation source of the target distribution and the forward-KL fitting approach, and states that KL regularization keeps importance weights well behaved by limiting the target distribution's deviation from the rollout policy.

The forward-KL objective is estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. Combining SNIS with forward KL allows actor fitting on actions sampled from the rollout policy, while KL regularization controls the target distribution's deviation from the rollout policy. The abstract explicitly names SNIS as the estimator and the rollout policy as the action source, and explains the role of KL regularization for importance weights.

Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains, and computational benchmarks show faster actor updates than REPPO. The abstract reports three kinds of results together: task performance, sample efficiency, and actor update speed, with the REPPO comparison belonging to computational benchmarks. The abstract states the experimental platforms and ablations, sample-efficiency gains, and the update-speed advantage over REPPO; specific numbers, task counts, and statistical details are not given in the abstract.

Perspective

The work targets online maximum-entropy reinforcement learning in continuous control, in settings that sample actions from a rollout policy and perform policy improvement using critic values; the abstract's validation covers MuJoCo Playground and ManiSkill, including ablations and a computational benchmark against REPPO. For researchers and practitioners who want to avoid differentiating the critic with respect to actions, or who care about the computational cost of actor updates, FERPO offers a reproducible algorithmic route: derive a target action distribution from a regularized policy-improvement objective, then fit the actor with a forward-KL objective estimated by SNIS.

The abstract does not give per-task performance numbers, the magnitude of sample-efficiency gains, specific ablation configurations, or the measurement conditions of the computational benchmark, so the robustness of the advantages across tasks and hyperparameters is difficult to judge. The behavior of forward KL and SNIS importance weights depends on KL regularization limiting the target distribution's deviation from the rollout policy, and how appropriate that limit is across settings remains an open question. The abstract also does not state the comparison scope beyond REPPO, nor transfer performance on tasks outside MuJoCo Playground and ManiSkill.

Sources