AnisoWM swaps isotropic regularization for a learnable diagonal covariance target and beats LeWM's planning success in all four visual control environments
Synopsis
The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
Interpretation
The paper shows that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost. Previously, SIGReg's isotropic Gaussian target was motivated by representation-quality criteria under linear and nonlinear probing, not by whether the resulting Euclidean distances suit planning; this work ties that target directly to planning rankings. In a linear Gaussian setting, as process noise vanishes the joint objective selects an approximately whitened representation, so Euclidean latent distance induces inverse-state-covariance weighting; even with exact conditional-mean prediction and vanishing prediction loss, a finite-horizon construction retains positive planning regret; in a two-dimensional MLP toy experiment, as the noise scale decreases, held-out conditional-mean prediction error drops substantially while mean normalized physical planning regret stays near the inverse-covariance reference across ten seeds.
It introduces AnisoWM with ΛReg, replacing the fixed isotropic Gaussian target with a learnable diagonal covariance whose variance allocation across latent directions is determined by predictive training under fixed-trace and condition-number constraints. Unlike changing the predictive structure (Fast-LeWM), adding multi-horizon and reachability supervision (RC-aux), or replacing/augmenting the terminal planning cost (TRM, Decision-Metric Alignment), this changes only the representation regularizer: the prediction objective, predictor architecture, and Euclidean planner stay the same, and the covariance target is used only during training and can be discarded at planning time. Theory gives the limiting metric selected under the learned target and the condition for proportionality to the task metric; a two-dimensional closed-form analysis shows that for a particular covariance ordering the limiting regret is zero at one intermediate anisotropy bound, lower than the isotropic baseline below it, and higher above it, so excessive anisotropy raises regret.
Across four visual control environments, AnisoWM improves planning success over LeWM in all four, and its latent costs agree better with task outcomes. Relative to the LeWM baseline, the gains come from representation geometry rather than the planner or predictor; the paper also reports ordering diagnostics under both a representation-only cost and a cost that includes predictor rollout, separating where a difference arises. The primary comparison uses 192-dimensional representations, a single shared anisotropy bound (at most a twofold ratio between target variances), and three training seeds per environment; success-rate gains appear in all four environments, ordering diagnostics improve in all four, with PushT nearly unchanged under the representation-only cost, indicating that the contributions of representation geometry and predictor rollout differ across environments; in local cost geometry, Spearman rank correlation between latent cost and task cost rises in both Cube and PushT.
Under the same anisotropy bound, the learned target spectra differ across environments, and loosening the anisotropy bound is not monotonically beneficial. This shows the constraint only defines a feasible family; the actual variance allocation is determined by predictive training and the training distribution rather than by the bound itself. Under the shared bound, the number of coordinates above the mean target variance ranges from 95 in PushT to 124 in Reacher; in TwoRoom the allocation keeps evolving after reaching the condition-number bound; across the anisotropy-bound sweep the preferred value differs by environment, and the sweep uses single training runs, which the authors frame as a sensitivity diagnostic rather than a tuned comparison.
Perspective
The result targets JEPA-style latent world models that plan to goals with Euclidean latent distance, especially settings that follow LeWM's data, architecture, and visual goal-planning protocol; the method replaces only the training-time regularization target, leaving the prediction objective, predictor architecture, and planner unchanged, so it can be layered onto an existing LeWM pipeline. The analysis covers metric selection and finite-horizon separation in a linear Gaussian setting and a two-dimensional closed-form regret as a function of the anisotropy bound; the experiments cover the four visual control environments TwoRoom, Reacher, PushT, and Cube, with a primary comparison of three seeds per environment under a shared anisotropy bound. For researchers who want to adjust representation geometry without touching the planner, this offers a reproducible route.
The anisotropy-bound sensitivity sweep uses single training runs, which the authors themselves frame as a sensitivity diagnostic, so the preferred bound per environment still needs confirmation with more repetitions; under the representation-only cost, ordering in PushT is nearly unchanged, indicating that the relative contributions of representation geometry and predictor rollout vary by environment and are not fully separated; the theoretical zero-regret and alignment conditions rely on a linear Gaussian model, specific covariance structures, and a two-dimensional closed-form setting, with extension to higher-dimensional nonlinear systems given as a stability argument rather than a selection theorem; in addition, this summary is based on the paper's full text and abstract without the underlying figure data, so specific numerical details should be checked against the original.
