DSReg provably recovers individual world latents without reconstruction, improving sparse control and editing across synthetic, pixel-trained, and external-renderer benchmarks
Synopsis
The work introduces DSReg (Dependency-Sparsity Regularization), which, on top of the linear-identifiability premise provided by LeJEPA, proves that under Structural Diversity—distinct dependency footprints of latents on observations—each world latent can be recovered up to signed permutation with no reconstruction, no decoder, and no labels; the method applies post hoc to any linearly identified representation, reuses trained checkpoints, and across synthetic regimes, world-model probes, pixel-trained encoders, and external renderers preserves dense prediction while improving individual-latent recovery and downstream sparse control, editing, and few-shot readout.
Interpretation
The paper proves that, given LeJEPA's linear-identifiability premise (in a Gaussian latent world the alignment optimum is unique up to an orthogonal transformation), Structural Diversity (pairwise-distinct dependency footprints of latents on observations) together with a faithfulness-type functional no-cancellation condition suffices for minimizing dependency-support sparsity to recover candidate latents up to signed permutation, so each estimated variable is one true world latent. Prior reconstruction-free, decoder-free, auxiliary-supervision-free routes such as JEPAs identified the latent state only up to a linear transformation, leaving individual latents mixed; the authors state this is the first component-wise identifiability result requiring no reconstruction, labels, or distributional asymmetry, and that Structural Diversity is strictly weaker than prior structural conditions. Stated as theorems (Theorem 2 and the approximate-recovery Theorem 3 under Jacobian perturbation) with proofs; Proposition 1 argues via implication chains on footprint sets that Structural Diversity is strictly weaker than Non-Inclusion, both forms of Structural Sparsity, and Structural Variability, with a witness counterexample.
The authors provide an actionable objective, DSReg: on an already linearly identified representation, search for an additional orthogonal rotation whose candidate latents have the sparsest dependency Jacobian support; it can be optimized jointly or post hoc, and the post-hoc fit reuses trained checkpoints with an ablation indicating no loss relative to joint training. Turns the identifiability theory into a sparsity criterion that reads only observations and estimated latents, requiring no decoder, observation likelihood, or reconstruction error, and gives a scalable implementation that factors the criterion through neighborhoods to avoid materializing large matrices. Provides a finite-sample relaxation (Equation 4) and three interchangeable evaluations (materialized, exactly factored, and an unbiased row-sampled estimator), reporting that on one GPU the per-step cost stays essentially flat as latent dimension scales to 2048 and observed dimension to 8192; a schedule comparison table shows joint and decoupled recovery are comparable.
Experiments show recovery succeeds exactly where Structural Diversity holds and fails where it does not: near-ceiling recovery in synthetic regimes with pairwise-distinct footprints (including nested and minimal-difference cases), while with identical footprints individual recovery is impossible from support alone yet the pair's span is still recovered; on pixel-trained encoders DSReg approaches the label-using Procrustes oracle while latent-only rotations such as PCA, Varimax, and FastICA leave latents mixed. Aligns the theory's predicted success/failure boundary with testable predictions and attributes the gain to the observation-side dependency signal rather than the optimizer; also shows the gain persists on external renderers (Gaussian 3DShapes and quarter-orientation dSprites). Five runs per point on synthetic orbits, five to fifteen random seeds for learned visual encoders and ten to twenty for the remaining families; reports unchanged dense readout (R²) while latent MCC, few-shot single-latent readout, and sparse-use scores improve; on dSprites reports one-sided Mann–Whitney tests against two VAE baselines.
Sparse modules (visual editing, sparse model predictive control, short-horizon rollout prediction, transition-surprise detection) perform better on recovered latents because such modules act through a few variables at a time; in a rotated basis the same capped model becomes systematically misspecified, whereas in the physical basis the dynamics are approximately sparse and the model is close to well specified. Grounds the question of why individual latents are preferable to a well-spanned mixture in concrete downstream tasks, and notes the probes do not use the permutation or signs left free by the theory, so they are invariant to exactly the ambiguity the theory does not resolve. Evaluated on six probes across TwoRoom, PushT, and FetchSlide, with thresholds and budgets shared between the two representations and the fitted rotation as the only difference; also reports dense readouts and retrieval costs unchanged between the two representations.
Perspective
The result targets representations that already satisfy the LeJEPA linear-identifiability premise (Gaussian latent world, stationary additive-noise transition, SIGReg Gaussian constraint) and that satisfy Structural Diversity; in that setting DSReg applies post hoc to any linearly identified representation, reuses trained checkpoints, and lets a controller, dynamics model, or monitor act through individual variables while dense readouts remain unchanged. The authors describe Structural Diversity as a mild requirement: factors with identical footprints cannot be distinguished from support alone and would need extra signals such as temporal structure or cheap interventions; they also state that current benchmarks are simulated or rendered and that DSReg has not yet been tested on a physical robot, which they name as the most exciting future work.
A careful reader would still watch: whether Structural Diversity holds broadly in real data, especially when global factors (illumination, camera gain, global style) create nested footprints that the theory still covers, while identical footprints would need extra signals; the Gaussian latent-world premise is a foothold rather than a ceiling, and how extensions of linear identifiability to richer worlds change the guarantee remains open; benchmarks are simulated or rendered, so behavior on a physical robot is unknown; and this material is full text in which figure and appendix-table values appear as placeholders, so exact numbers should be checked against the original figures.
