Skip to main content
Back to timeline
arXivSource publication:

FOCUS uses KL-guided modality switching to transfer privileged-state training to an RGB-D policy, raising average test success from 0.71 to 0.93 across five IsaacLab manipulation tasks

Synopsis

FOCUS is a single-stage PPO framework that keeps the critic on privileged state while an actor switches its rollout source between privileged-state and RGB-D latents according to the KL divergence between the two modalities' action distributions, combined with representation alignment; across five IsaacLab manipulation tasks it raises average test success from 0.71 to 0.93 and end-to-end budget-normalized training-success AUC from 0.47 to 0.65 relative to the strongest RGB-D-at-test baseline on each task.

Source-provided article image: FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment
Figure 1 ·

Figure 1: Overview of FOCUS. During training, cross-modal policy disagreement is aggregated into running statistics that determine the RGB-D exposure probability λ r \lambda_{r} for subsequent actor rollouts. A KL-guided rollout switch samples either the privileged-state or RGB-D pathway and keeps the selected pathway fixed throughout the rollout, while the critic remains anchored to privileged state. At inference, privileged state and the switching mechanism are removed, and the actor operates only from RGB-D, proprioception, and the previous action.

arXiv

Interpretation

FOCUS turns privileged-to-RGB-D transfer into single-stage PPO training: the critic is anchored to privileged state for value estimation, while a KL-guided switch selects whether the actor's rollout uses the privileged-state latent or the RGB-D latent, with the modality held fixed within each rollout. Unlike teacher-student or distillation pipelines that separate privileged teacher training from visual student transfer, FOCUS performs the transition online within one PPO run and uses cross-modal policy disagreement to decide which observation pathway generates on-policy experience. The method is fully formulated in Section 3, the training algorithm is given in Appendix A, Appendix B provides an exposure-weighted bound analysis of the KL-guided switch, and Appendix K lists implementation defaults.

KL-guided modality switching sets the RGB-D rollout probability from the KL divergence between the same policy's action distributions under privileged-state and RGB-D inputs, and an upper-quantile KL budget lowers RGB-D exposure while rare high-disagreement states remain. Relative to a fixed ramp or a mean-KL-only variant, the upper-quantile budget suppresses exposure to rare high-disagreement states; ablations report a fixed ramp reaching normalized AUC 0.361 versus FOCUS's 0.612 on Pick-and-Place and test success 0.822 versus 0.850 on Push-Cube. Ablations are reported on the three more discriminative tasks; Proposition 1 in Appendix B gives mean and quantile exposure-weighted mismatch bounds and notes that the implemented controller approximates them using lagged running statistics.

Representation alignment reduces cross-modal policy disagreement by freezing the actor and fusion module and updating only the RGB-D encoder, so matched privileged-state and RGB-D inputs induce similar action distributions. This policy-output alignment term corresponds to the directional disagreement monitored by the KL switch, tying representation learning directly to action selection rather than to latent similarity alone. Appendix C gives the full alignment objective decomposition (aligned-target regularization, raw-latent regularization, policy-conditioned alignment, optional decoding); Lemma 1 in Appendix B.4 bounds KL sensitivity to fused-input mismatch under a shared-covariance Gaussian policy and a local Lipschitz assumption.

Across five IsaacLab manipulation tasks, FOCUS raises average test success from 0.71 to 0.93 and end-to-end budget-normalized training-success AUC from 0.47 to 0.65 relative to the strongest RGB-D-at-test baseline on each task; on Pick-and-Place test success rises from 0.47 to 0.86 and AUC from 0.12 to 0.61. AUC is computed over the complete training procedure: FOCUS counts both privileged-state and RGB-D rollouts, while teacher-student counts teacher PPO, RGB-D distillation, and student PPO fine-tuning, so the comparison covers end-to-end budget rather than only the final student stage. Each method is trained with two training seeds and evaluated with five evaluation seeds per checkpoint, with test performance measured as success rate over a fixed 9000-step evaluation rollout; the authors note states-only is a privileged-observability reference rather than an upper bound and shows substantial variation across its two training runs.

Perspective

The work targets PPO manipulation settings where privileged state is available in simulation and only RGB-D, proprioception, and the previous action are available at test time, on IsaacLab tasks Reach-Cube, Lift-Cube, Push-Cube, Open-Drawer, and two-block Pick-and-Place, with simulation at 100Hz, control at 50Hz, and 5-second episodes. It addresses the privileged-state-to-RGB-D observation gap and is orthogonal to the sim-to-real gap; appearance shift, sensing noise, calibration error, latency, and dynamics mismatch would still require dedicated sim-to-real methods or real-world fine-tuning. The authors recommend evaluating on physical robots, broader manipulation suites, and less engineered training signals, and using more extensive task and seed coverage to isolate when the auxiliary alignment components are necessary versus when the lightweight FOCUS-simple core is sufficient.

Evaluation is limited to simulation, relatively short-horizon manipulation, and two training seeds, and the authors explicitly interpret ablations as component-level trends rather than universal claims about which auxiliary losses are necessary; some runs converge to locally successful but suboptimal behaviors, and the method is sensitive to task-specific reward engineering. The states-only reference shows substantial variation on Lift-Cube, and the authors do not read FOCUS's higher mean success there as evidence that RGB-D is intrinsically superior to privileged state; the asymmetric actor-critic baseline is unstable on some tasks and the present evaluation does not isolate the source. FOCUS-simple matches or exceeds the full version on two of three discriminative tasks, so which auxiliary alignment terms are necessary under which tasks remains open. In addition, some ablation values in the main-text table are not fully rendered in the loaded text, so item-by-item comparison should consult the original table.

Sources