Skip to main content
Back to timeline
arXivSource publication:

CTWM restores commute-time-faithful embeddings via a log-determinant regularizer, matching or beating LeWM on continuous goal-reaching benchmarks with half the parameters

Synopsis

The work shows that existing self-supervised methods, which commonly enforce isotropic representations to prevent collapse, tend to degrade the eigenvalue-dependent scaling required by the graph Laplacian spectral embedding and thus misrepresent commute times; it introduces Commute-Time-Preserving World Models (CTWMs), combining a latent displacement predictor with a log-determinant regularizer, which provably recover the correctly scaled Laplacian representation under reversible deterministic dynamics and at the predictor's fixed point, and in numerical simulations matches or outperforms LeWM, a task-agnostic baseline, on several complex continuous goal-reaching benchmarks while using half the parameters.

Source-provided article image: Learning Commute-Time-Preserving World Models for Planning
Figure 1 ·

Figure 1: (a) Overview of CTWM: a shared encoder ϕ \phi maps observations s s and s ′ s^{\prime} to the representations z z and z ′ z^{\prime} respectively. The transition model T T conditioned on action a a , predicts a residual displacement Δ \Delta that is added to the current embedding via a residual connection. The loss ℒ ⁡ ( ϕ , T ) \mathcal{L}(\phi,T) is computed from the predicted and actual next-state embeddings. (b) Geometric interpretation of the loss ℒ ⁡ ( ϕ , T ⋆ ) \mathcal{L}(\phi,T^{\star}) in the latent space; arrow colors refer to terms in equation 6 . The loss terms encourage predictable transitions (teal), penalize large distances between consecutive points (blue), and prevent collapse of representation (magenta). (c) Latent distances reflect commute-time distance rather than spatial proximity. Left: Two-room connected by a door; s 1 s_{1} and s 3 s_{3} are adjacent but separated by a wall (dashed line). Right: in the CTWM embedding the rooms unfold around the door, placing s 1 s_{1} near s 2 s_{2} and far from s 3 s_{3} .

arXiv

Interpretation

The paper identifies a concrete failure mode: isotropic representation constraints, widely used to prevent representational collapse, weaken the eigenvalue-dependent scaling required by spectral embeddings, so latent distances no longer accurately correspond to commute times in the environment. Whereas prior work treats isotropy as a general anti-collapse device, this work places that constraint in direct tension with commute-time fidelity. The claim comes from the paper's analysis and demonstration of existing methods, stated in the abstract as 'here we show', without reported experimental scale or numbers.

It proposes CTWM, replacing isotropic constraints with a latent displacement predictor plus a log-determinant regularizer, preserving correct eigenvalue scaling while avoiding collapse. Unlike existing self-supervised routes that rely on isotropy for anti-collapse, CTWM uses a log-determinant term and provides a provable recovery result at the predictor's fixed point under reversible deterministic dynamics. The abstract states a theoretical guarantee ('provably recover'), explicitly conditioned on reversible deterministic dynamics and the fixed point.

In numerical simulations on several complex continuous goal-reaching benchmarks, CTWM matches or outperforms LeWM, a task-agnostic baseline, while using half the parameters. Relative to the task-agnostic LeWM baseline, CTWM retains performance while substantially reducing parameter count. Evidence is a benchmark comparison in numerical simulations; the abstract reports no task counts, metric values, or statistical tests.

Perspective

The work targets agents that plan in latent space for goal reaching in large continuous environments: when instantiating the graph Laplacian is intractable, CTWM offers a self-supervised route to commute-time-faithful embeddings. Its theoretical recovery result applies under reversible deterministic dynamics and holds at the predictor's fixed point; the numerical simulations cover several complex continuous goal-reaching benchmarks and compare against the task-agnostic LeWM baseline. For researchers and practitioners seeking comparable planning performance with fewer parameters, this method provides a reproducible alternative.

Readers may still watch: whether the log-determinant regularizer preserves commute-time scaling under non-reversible or stochastic dynamics; how the fixed-point assumption is satisfied and checked in practice; the magnitude of 'matches or outperforms', task coverage, and variance in the numerical simulations; and whether halving parameters changes training cost or convergence speed. Because only the abstract is available here, figures, ablations, and statistical details in the full text are not yet incorporated, so these remain open directions.

Sources