Skip to main content
Back to timeline
arXivSource publication:

D-JEPA supervises relations among candidate futures with executed outcomes, reaching 87.89% on PushT and a 17-point gain on physical robots

Synopsis

D-JEPA introduces a decision-aligned latent world model: it first quantifies a decision-local prediction gap in which, among the few futures competing for execution, a candidate predicted closer to the goal can fail while an available alternative succeeds; it then uses executed outcomes to supervise a bounded, permutation-equivariant set operator that learns decision-relevant relations among candidate futures, and realizes the learned decision structure in JEPA-compatible future representations so aligned actions can be read out through native latent-distance planning; it reaches 87.89% success on PushT, a 15.04-point average gain on RoboTwin, a 17-point gain on physical robot tasks, and raises mean driving PDMS from 57.36 to 95.34.

AI-generated editorial illustration: D-JEPA: A Decision-Aligned Latent World Model

Interpretation

The paper identifies and quantifies a decision-local prediction gap: predictive geometry remains globally informative yet weakly distinguishes the candidates closest to execution. Prior work on latent world models focused on predicting futures and on plan-cost fidelity; this work reformulates the problem as how executed outcomes should supervise relations among candidates that share the same observed context and goal. Across a 96-start PushT audit, each model prefers a failing candidate on eight starts despite an available successful alternative; among mixed-outcome top-four shortlists, mean inversion rates reach 49.0% for LeWM and 38.2% for TD-JEPA; within the four lowest-cost candidates, within-start Spearman correlation falls below 0.13 (LeWM 0.90 to 0.11, TD-JEPA 0.80 to 0.13).

D-JEPA learns decision-relevant relations among candidate futures with a bounded, permutation-equivariant set operator, complemented by restricted predictor adaptation and a shared ordinal interface across predictive geometries. Unlike replacing the prediction objective or applying only scalar confidence gating, the method retains complete-set predictive context while concentrating supervision on the decision-relevant subset and bounding changes to the underlying representation and decision score. In matched ablations, relational alignment is the strongest individual intervention (TD-JEPA 76.95% to relational alignment 87.11% to calibrated composition 87.89%); on the 128-start mechanism population the same relational operator with dual LeWM and TD-JEPA evidence reaches 93.75%, versus 84.38% and 83.59% from either source alone; the four-geometry instance matches calibrated four-source fusion.

The learned decision structure can be written into JEPA-compatible future representations, so aligned action choices are recovered directly through the native latent-distance planning interface. This makes decision alignment more than a training-time re-ranking: it is expressed in the future geometry available at deployment, preserving the native planning interface. After exact ordinal realization, native goal distance recovers 100% of relational action choices and candidate ranks; relational readout gives 87.11% success at 35.03 ms, calibrated composition 87.89% at 52.16 ms, and ordinal realization 87.11% at 34.02 ms; on Reacher, temporal transport learns signed updates across five future steps within the prescribed bound, with an observed maximum per-step L2 update of 0.0017227774.

Decision alignment improves executed choices across distinct physical dynamics, pretrained action-producing models, physical robots and driving scenes, and retains gains under shape and visual changes. Evaluation spans latent control, manipulation, pretrained action models, physical robots and driving, with comparisons sharing starts, goals, candidate actions and execution horizons while retaining the reference planner and search budget. PushT 87.89% and Reacher 93.75%; RoboTwin averages 61.72% to 76.76% across four tasks (+15.04 points); two physical PiPER tasks rise from 64.0% to 81.0% (+17.0 points, with 12/3 gains-to-losses on PushT and 10/2 on stacking); mean PDMS across seven driving scenes rises from 57.36 to 95.34; improvements appear on all three held-out PushObj geometries and all seven appearance conditions, with object-colour change yielding the largest appearance gain.

Perspective

This work targets the decision stage of choosing which available future to execute, and applies to latent planners, pretrained VLA action chunks, Drive-JEPA trajectories and V-JEPA 2-AC candidates; it does not change the underlying predictor, search budget or low-level controller. The method improves as candidate availability grows (81.96% with eight candidates to 87.11% with 63), so it suits settings that can supply larger candidate pools. Shape evaluation uses 100 held-out starts per geometry and visual evaluation uses 50 fresh starts per condition, both selecting from 63 candidates and executing the complete 25-control sequence; driving reports an equal-weight mean of seven scene-level PDMS values.

The decision-local gap diagnosis rests on 96 PushT starts and recorded executions, so how general it is across task families still needs more evidence; the gains from relational alignment and predictor adaptation depend on calibrated gates and thresholds chosen on calibration subsets, which would need recalibration for new tasks. The ordinal interface's scale invariance concerns ordinal evidence, while dense descriptors retain each model's own geometry, so cross-geometry integration depends on the quality of each source geometry. Driving results report an equal-weight mean over seven scenes drawn from three source logs, a limited sample. Physical-robot results rest on 50 paired trials per task while retaining the same V-JEPA 2-AC planner and search budget. In addition, this is a full-text parse, but some figures and appendix details appear as textual descriptions, so reproducing specific numbers still requires consulting the original appendices.

Sources