Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

BLS refines task vectors at test time, letting zero-shot policies surpass OLS and online adaptation baselines on OGBench, DMC, and HumEnv

The authors propose BLS, an offline zero-shot task-inference method for Behavioral Foundation Models that starts from an ordinary-least-squares (OLS) task vector and refines it with a soft-margin contrastive loss plus a trust-region loss to reduce both reward reconstruction error and successor-measure mismatch; they prove a suboptimality gap upper bound characterized by successor-measure and reward residuals, and empirically BLS outperforms OLS, ZOL, ReLA, and LoLA on average across five feature representations on OGBench (proprioceptive and pixel observations), 20 DMC tasks, and 45 HumEnv humanoid control tasks, with negligible inference overhead.