Public articles linked to the same research event.
arXiv The work introduces TwinJEPA, which augments JEPA-based offline zero-shot control by identifying approximately matched states across offline trajectories and constructing preference pairs via goal-conditioned reward relabeling, then learning two training-only objectives, reward-gap regression and preference classification, to obtain action-preferred predictive representations; on long-horizon navigation and continuous-control benchmarks with state- and pixel-based observations, it yields positive benchmark-level mean differences across all matched state-based evaluations, with analyses indicating larger gains tend to arise when local action alternatives provide more informative outcome contrasts.
The work introduces TwinJEPA, which augments JEPA-based offline zero-shot control by identifying approximately matched states across offline trajectories and constructing preference pairs via goal-conditioned reward relabeling, then learning two training-only objectives, reward-gap regression and preference classification, to obtain action-preferred predictive representations; on long-horizon navigation and continuous-control benchmarks with state- and pixel-based observations, it yields positive benchmark-level mean differences across all matched state-based evaluations, with analyses indicating larger gains tend to arise when local action alternatives provide more informative outcome contrasts.
The work introduces TwinJEPA, which augments JEPA-based offline zero-shot control by identifying approximately matched states across offline trajectories and constructing preference pairs via goal-conditioned reward relabeling, then learning two training-only objectives, reward-gap regression and preference classification, to obtain action-preferred predictive representations; on long-horizon navigation and continuous-control benchmarks with state- and pixel-based observations, it yields positive benchmark-level mean differences across all matched state-based evaluations, with analyses indicating larger gains tend to arise when local action alternatives provide more informative outcome contrasts.
The work introduces TwinJEPA, which augments JEPA-based offline zero-shot control by identifying approximately matched states across offline trajectories and constructing preference pairs via goal-conditioned reward relabeling, then learning two training-only objectives, reward-gap regression and preference classification, to obtain action-preferred predictive representations; on long-horizon navigation and continuous-control benchmarks with state- and pixel-based observations, it yields positive benchmark-level mean differences across all matched state-based evaluations, with analyses indicating larger gains tend to arise when local action alternatives provide more informative outcome contrasts.