Skip to main content
Back to timeline
arXivSource publication:

BLS refines task vectors at test time, letting zero-shot policies surpass OLS and online adaptation baselines on OGBench, DMC, and HumEnv

Related research and updates

Synopsis

The authors propose BLS, an offline zero-shot task-inference method for Behavioral Foundation Models that starts from an ordinary-least-squares (OLS) task vector and refines it with a soft-margin contrastive loss plus a trust-region loss to reduce both reward reconstruction error and successor-measure mismatch; they prove a suboptimality gap upper bound characterized by successor-measure and reward residuals, and empirically BLS outperforms OLS, ZOL, ReLA, and LoLA on average across five feature representations on OGBench (proprioceptive and pixel observations), 20 DMC tasks, and 45 HumEnv humanoid control tasks, with negligible inference overhead.

Source-provided article image: Task Inference Beyond Least Squares in Behavioral Foundation Models
Figure 1 ·

Figure 1: Ordinary least squares (OLS) aims to minimize the total errors of reconstructed rewards. Its performance highly depends on the quality of state representations, the state coverage within the dataset, and the distribution of rewards. Because OLS can be biased toward the majority of states, low-reward states ( N N ) may be assigned with high-value rewards and vice versa for high-reward states ( G G ). The successor measure of the retrieved zero-shot policy can therefore diverge from that of the optimal policy. BLS instead adopts an objective that reduces the reward reconstruction errors while enlarging the margin between high- and low-reward states. Thus, the successor measure of the retrieved zero-shot policy is encouraged to stay closer to the optimal policy’s.

arXiv

Interpretation

The paper derives an upper bound on the suboptimality gap of a zero-shot retrieved policy that depends jointly on the reward reconstruction residual and the policy-averaged successor-measure residual, and notes that OLS optimizes only the former. Prior BFM task inference typically targets OLS reward reconstruction error alone, leaving successor-measure drift outside the inference objective; this work places both residuals in a single bound. Theorem 4.1 with its proof in Appendix A, together with the analysis that OLS pursues only the reward-agreement condition.

It proposes BLS: anchored at the OLS solution, a soft-margin contrastive loss enlarges the margin between reconstructed rewards of high-reward and low-reward states, while a trust-region loss keeps the task vector close to the anchor, optimized by gradient descent. Unlike ReLA and LoLA, which need task-specific online interaction, BLS refines the task vector offline at test time directly from reward-feature samples; unlike ZOL's distribution correction, BLS leaves the representation, successor measure, and policy training untouched. Algorithm 1 and the objectives in Section 4.2; Appendix B specifies inference samples, Adam steps and learning rate, unit-sphere projection, and reward normalization.

On OGBench, BLS improves the average performance of all five feature representations under both observation modalities, with particularly large gains for HILP and TD-JEPA under proprioceptive observations. No prior method validated task-inference improvements across so many differently pretrained feature representations at once. Table 1 reports success rates on nine sparse-reward navigation and manipulation tasks for Laplacian, ICVF, HILP, FB, and TD-JEPA, with means and standard deviations over three random seeds.

BLS surpasses online adaptation methods in the offline inference setting and remains effective on dense-reward tasks: it improves 18 of 20 DMC tasks and 42 of 45 HumEnv tasks. ReLA and LoLA require reward-labeled interaction episodes per task, whereas BLS needs no task-specific interaction; BLS also refines strong anchors beyond OLS, such as the FB-CPR goal-embedding anchor on HumEnv. Cross-representation average success rates in Table 3, DMC and HumEnv results in Figure 3, and HumEnv goal-reaching and five FB-CPR checkpoint results in Tables 8 and 9.

Perspective

The method targets test-time use of an already pretrained BFM: the representation, successor measure, and policy stay frozen while only the task vector is updated, and it applies when reward-feature samples can be drawn from an offline replay buffer. For sparse-reward tasks a small weight on the soft-margin contrastive loss is preferable; for dense-reward tasks staying close to the OLS anchor is preferable, and the paper parameterizes trust-region strength via a cosine-similarity target. It can also refine strong anchors other than OLS, such as the task vector formed from the goal state's backward embedding in HumEnv goal-reaching, and weighted reward inference as an anchor.

The soft-margin objective is explicitly described by the authors as an empirical surrogate for successor-measure mismatch rather than a theoretical guarantee, and its weight against the anchor loss varies with reward sparsity, so it must be chosen per task type. Constructing the high-reward set on dense-reward tasks relies on a reward threshold; the paper reports threshold sensitivity, but jointly selecting threshold and weight remains open. Results are reported as success rates and returns on simulated benchmarks, and transfer to real robots or broader task distributions is not verified in the text.

Sources