VIGOR uses latent-space consistency to give model-based RL zero-shot generalization to unseen visual distractions, beating the second-best baseline by 3.4% on DMC and 43.6% on Robosuite
Related research and updatesSynopsis
The authors propose VIGOR, a framework combining asymmetric weak-to-strong augmentation, dynamics-level consistency via direct latent regression that enforces augmentation-invariant transition predictions, and encoder-level stabilization, achieving zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its model-based RL backbone; it outperforms state-of-the-art model-free and model-based baselines on the DeepMind Control Suite and Robosuite, surpassing the second-best baseline by 3.4% and 43.6% respectively, and ablations show that replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, indicating that latent-space consistency rather than the augmentation choice drives robustness.
Figure 1: Challenges of unseen visual distractions. Left: During training (top), clean observations s t train s_{t}^{\text{train}} are encoded into in-distribution latents z t ∈ 𝒵 train z_{t}\in\mathcal{Z}^{\text{train}} , where the dynamics model d θ d_{\theta} produces accurate transitions. At test time (bottom), visual distractions yield out-of-distribution latents z t ′ ∉ 𝒵 train z^{\prime}_{t}\notin\mathcal{Z}^{\text{train}} , causing erroneous transitions. Right: These deviations compound over horizon H H , leading to divergent rollouts and planning failures. The proposed VIGOR framework targets this failure mode by aligning perturbed rollouts ( z t ′ , z t + 1 ′ , … , z t + H ′ ) (z^{\prime}_{t},z^{\prime}_{t+1},\dots,z^{\prime}_{t+H}) with their in-distribution counterparts ( z t , z t + 1 , … , z t + H ) (z_{t},z_{t+1},\dots,z_{t+H}) , anchoring predictions to the task-relevant manifold 𝒵 train \mathcal{Z}^{\text{train}} (mechanism described in Section 3 ).
arXivInterpretation
The paper identifies a two-level vulnerability in model-based RL: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon, unlike model-free RL where encoder perturbations affect only single-step predictions. It reframes visual generalization from single-step encoder robustness to error accumulation in latent dynamics under recursive planning, offering a new problem characterization for method design. The claim is presented as a mechanistic contrast between model-free and model-based methods, serving as the paper's motivation; the abstract provides no quantitative validation.
VIGOR integrates three interdependent components: asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. It moves the augmentation-invariance constraint from the encoder output level down to the latent dynamics transition level, and adds encoder stabilization to counter the drift induced by that cross-augmentation supervision, forming a coupled three-part design. The abstract describes the method component-wise and gives no separate ablation numbers for each component; component necessity rests on the stated interdependence.
Evaluations on the DeepMind Control Suite and Robosuite show VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. It delivers a zero-shot visual generalization advantage while retaining the sample efficiency of the MBRL backbone, with markedly different margins across the two benchmarks. The abstract reports relative improvements on two benchmarks but omits absolute returns, number of random seeds, confidence intervals, and statistical tests.
Ablations show VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness. It shifts the attribution of effectiveness from a specific augmentation strategy to the latent-space consistency mechanism itself, reducing dependence on a particular augmentation design. The abstract states this as an ablation conclusion without listing the substituted augmentation families, the number of ablations, or corresponding performance values.
Perspective
The work targets model-based RL that plans within learned latent dynamics, aimed at settings with unseen visual distractions such as background variations, lighting changes, or camera shifts, with the goal of zero-shot generalization while retaining MBRL sample efficiency. Its validation scope is two benchmarks, the DeepMind Control Suite and Robosuite, compared against state-of-the-art model-free and model-based baselines. For readers, this means that if your system is a latent-planning MBRL setup deployed where visual distribution shift occurs, VIGOR's component split (augmentation pairing, dynamics-level consistency, encoder-level stabilization) can serve as a directly reusable design template; the demonstrated augmentation-agnosticism also suggests implementations need not be tied to a specific augmentation operator.
The abstract gives no absolute return values, number of random seeds, confidence intervals, or significance tests, so the comparability of the 3.4% and 43.6% relative gains across benchmarks still needs to be judged against the main tables. The ablation only states that substituting augmentations from distinct perturbation families preserves strong generalization, without listing the specific families or their performance, so the boundary conditions of augmentation-agnosticism await confirmation in the main text. In addition, evaluation is limited to two simulated benchmarks, DMC and Robosuite; behavior under real robotic visual distractions, and the independent contribution and degree of interdependence of the three components, are directions readers can continue to watch.
