Public articles linked to the same research event.
arXiv The authors propose BLS, an offline zero-shot task-inference method for Behavioral Foundation Models that starts from an ordinary-least-squares (OLS) task vector and refines it with a soft-margin contrastive loss plus a trust-region loss to reduce both reward reconstruction error and successor-measure mismatch; they prove a suboptimality gap upper bound characterized by successor-measure and reward residuals, and empirically BLS outperforms OLS, ZOL, ReLA, and LoLA on average across five feature representations on OGBench (proprioceptive and pixel observations), 20 DMC tasks, and 45 HumEnv humanoid control tasks, with negligible inference overhead.
The authors propose BLS, an offline zero-shot task-inference method for Behavioral Foundation Models that starts from an ordinary-least-squares (OLS) task vector and refines it with a soft-margin contrastive loss plus a trust-region loss to reduce both reward reconstruction error and successor-measure mismatch; they prove a suboptimality gap upper bound characterized by successor-measure and reward residuals, and empirically BLS outperforms OLS, ZOL, ReLA, and LoLA on average across five feature representations on OGBench (proprioceptive and pixel observations), 20 DMC tasks, and 45 HumEnv humanoid control tasks, with negligible inference overhead.
The authors propose BLS, an offline zero-shot task-inference method for Behavioral Foundation Models that starts from an ordinary-least-squares (OLS) task vector and refines it with a soft-margin contrastive loss plus a trust-region loss to reduce both reward reconstruction error and successor-measure mismatch; they prove a suboptimality gap upper bound characterized by successor-measure and reward residuals, and empirically BLS outperforms OLS, ZOL, ReLA, and LoLA on average across five feature representations on OGBench (proprioceptive and pixel observations), 20 DMC tasks, and 45 HumEnv humanoid control tasks, with negligible inference overhead.
The authors propose BLS, an offline zero-shot task-inference method for Behavioral Foundation Models that starts from an ordinary-least-squares (OLS) task vector and refines it with a soft-margin contrastive loss plus a trust-region loss to reduce both reward reconstruction error and successor-measure mismatch; they prove a suboptimality gap upper bound characterized by successor-measure and reward residuals, and empirically BLS outperforms OLS, ZOL, ReLA, and LoLA on average across five feature representations on OGBench (proprioceptive and pixel observations), 20 DMC tasks, and 45 HumEnv humanoid control tasks, with negligible inference overhead.