CrossBFM distills Unitree G1's latent behavior space onto three humanoids in under one GPU-hour, losing only 0.025 rad in tracking
Synopsis
CrossBFM freezes the backward map of a Behavior Foundation Model pretrained on the Unitree G1 and, using the frame-level cross-embodiment correspondence supplied by retargeted data, distills that latent space onto three humanoids (Inhouse M3, Booster T1, Fourier N1) with a unified encoder that has no robot-specific parameters, training the encoder in under one GPU-hour and each tracker in about 10 more GPU-hours; all three prompting modes transfer (motion tracking, goal reaching between poses, reward optimization), the latent-conditioned policy loses only 0.
Interpretation
A unified encoder with no robot-specific parameters distills the frozen source BFM's backward map onto multiple embodiments, handling all training robots in a single training run. Prior frameworks such as BFM-Zero and UFO train once per robot at more than 100 GPU-hours each, and a second run produces a latent space unrelated to the first; CrossBFM treats the latent space itself as the transferable asset, using a fixed-width input (8 key-body indices, 33 canonical joint slots, missing slots masked and zero-padded) so robots with different degrees of freedom and topologies share one network. The paper reports encoder training in under one hour on a single consumer GPU and trackers trained with conventional PPO in about 10 GPU-hours; the appendix gives roughly M parameters for the unified encoder versus roughly M each for per-robot encoders, and reports higher latent agreement among the three new robots for the unified encoder than for separately trained per-robot encoders.
Retargeted data serve as a correspondence oracle, reducing target-side latent learning to supervised regression onto the frozen source latent. Earlier cross-embodiment work either conditions one policy on an embodiment description for a single task or aligns state spaces explicitly with optimal transport or graph matching; here frame-aligned retargeted motions on a shared timeline let the retargeting algorithm carry the correspondence, removing the simulator, RL training, and adversarial objective from the target side. The objective is frame-wise cosine regression onto the frozen source latent, with a pre-LayerNorm transformer encoder and a causal attention mask; appendix perturbation experiments show timing offsets are recovered exactly by a one-dimensional phase search, non-systematic Gaussian noise is tolerable up to about 0.1 rad at inference and 0.2 rad at training, and systematic bias is harmless up to roughly 0.1 rad but destructive by 0.3 rad.
Latent-conditioned trackers cover all three BFM prompting modes on the three distilled humanoids with performance close to joint-conditioned policies. Joint-conditioned policies see the exact joint angles at every step and are the natural upper bound for tracking; latent-conditioned policies receive only a single behavior latent yet still track motions, reach goal poses smoothly with no falls, and optimize all 41 reward prompts. The paper reports a gap of only 0.025 rad between latent- and joint-conditioned tracking; goal reaching is reported only for latent-conditioned policies because commanding discontinuous joint targets produces large torques and early termination; in reward optimization, using the source robot's true backward map is worse on M3 and N1 than CrossBFM, because the reward-weighted projection depends on which states the executing robot can actually reach.
Distillation scales on both data and robot count: a quarter of the corpus suffices, and unseen morphologically similar robots transfer partially. This gives an empirical path to lowering distillation cost and indicates that cross-embodiment generalization is closer to interpolation than extrapolation: a new robot does better when it has a close relative in the training set. Training encoders on randomly selected continuous segments totaling a quarter of the LAFAN corpus raises joint MAE from about 0.05 rad at full data to about 0.06 rad, costing only 5% of tracking performance, with breakdown only around 10% of the data; in leave-one-robot-out training, an encoder that never saw M3 recovers about 89% of tracking performance, while unseen T1 recovers only about 60%, which the paper attributes to T1's larger morphological difference and lack of a close relative in training.
Perspective
The result targets humanoid platforms that already have a frozen source BFM and frame-aligned retargeted motions: encoder training takes under one GPU-hour, each robot's latent-conditioned tracker about 10 GPU-hours, and adding a new robot requires only configuration such as a name map and a scale constant rather than extra training or architecture changes. It lets one latent vector command multiple embodiments at once and supports sampling behaviors inside the shared space with a flow-based generator, widening prompts from joint-space references toward behavior modes and text. The paper lists richer training data for a finer-grained behavior space and a larger robot training set for better generalization to unseen robots as future directions.
Several table values in the paper appear blank or as placeholders in the loaded text, so the specific numbers for tracking MAE, goal-reaching MAE, reward-optimization comparisons, the data-scaling curve, and the leave-one-robot table can only be relayed from the abstract and prose rather than checked cell by cell. Cross-embodiment generalization is currently demonstrated on morphologically similar robots, and the paper notes latent similarity drops when an unseen robot has no close relative in training, so extending generalization to more divergent embodiments and scaling the behavior space to finer granularity and longer-horizon loco-manipulation remain open directions. Real-robot validation covers the M3 and T1 robots, and the paper mentions small sim-to-real degradation whose magnitude across other hardware and tasks is worth tracking.
