TwinJEPA adds offline-mined action-preference supervision to JEPA control, yielding positive benchmark-level mean differences across all matched state-based evaluations
Related research and updatesSynopsis
The work introduces TwinJEPA, which augments JEPA-based offline zero-shot control by identifying approximately matched states across offline trajectories and constructing preference pairs via goal-conditioned reward relabeling, then learning two training-only objectives, reward-gap regression and preference classification, to obtain action-preferred predictive representations; on long-horizon navigation and continuous-control benchmarks with state- and pixel-based observations, it yields positive benchmark-level mean differences across all matched state-based evaluations, with analyses indicating larger gains tend to arise when local action alternatives provide more informative outcome contrasts.
Figure 1: Overview and key insight of TwinJEPA. A: A transition-wise temporal-prediction objective such as TD-JEPA learns goal-conditioned temporal structure from action-conditioned transitions, but does not explicitly compare alternative actions in similar states. B: TwinJEPA mines action-diverse transition pairs from local state neighborhoods and derives reward-gap and preference targets through goal-conditioned relabeling. C: Temporal prediction and pairwise supervision jointly train the shared representation to preserve trajectory-level temporal structure while encoding local distinctions between actions associated with different goal-conditioned outcomes. The auxiliary heads are used only during training and introduce no additional inference-time computation.
arXivInterpretation
Introduces TwinJEPA, which augments JEPA-based control with offline-mined action-preference supervision so that representations retain both long-horizon temporal structure and fine-grained distinctions among locally observed actions. Prior action-conditioned temporal prediction supervises actions in isolation at the transition level and provides limited information about which actions are preferable when similar states admit different goal-conditioned outcomes; TwinJEPA supplies that information through preference pairs. At the abstract level the paper states the method composition and motivation and reports evaluation on long-horizon navigation and continuous-control benchmarks covering state- and pixel-based observations.
Preference pairs are constructed by identifying approximately matched states across offline trajectories and relabeling with goal-conditioned rewards. The data source for preference learning is mined inside existing offline behavioral data rather than relying on additional interaction or online sampling. The abstract describes this construction explicitly but does not give the matching criterion's thresholds or data scale.
Training uses two complementary objectives: reward-gap regression preserves the magnitude of outcome differences, and preference classification captures the relative ordering of alternative actions. Magnitude and ordering are modeled separately rather than through a single preference signal. The abstract states that both objectives are used only during training and incur no additional inference-time cost.
Positive benchmark-level mean differences are obtained across all matched state-based evaluations, and analyses across domains and observation modalities indicate larger gains tend to arise when local action alternatives provide more informative outcome contrasts. Gain size is linked to the informativeness of local action contrasts rather than only reporting overall improvement. The abstract reports benchmark-level mean differences and cross-domain, cross-modality trend analyses, without specific values, sample sizes, or per-task results.
Perspective
The work targets offline zero-shot control: representations are learned from fixed behavioral data, and evaluation centers on long-horizon navigation and continuous-control benchmarks covering state- and pixel-based observations. Preference supervision is used only during training, so deployment adds no inference cost, making it suited to settings that already have offline trajectories and want to improve goal-conditioned control without extra online interaction. For researchers and practitioners who want to attach local action-comparison signals to JEPA-style representation learning, the pipeline of constructing preference pairs with two objectives can serve directly as a starting point.
The abstract reports benchmark-level mean differences and cross-domain, cross-modality trends without specific values, sample sizes, per-task results, or statistical tests, so the magnitude and stability of gains still require the main text. How the approximately matched states are judged, and how the number and quality of preference pairs affect results, are not expanded in the abstract. The weighting and interaction between the two training objectives, and behavior under different offline data coverage, are also directions a reader would keep asking about.
