Skip to main content
Back to timeline
arXivSource publication:

Vela replaces fixed-rate action chunks with continuous spline trajectories, reaching 45.3% average success on LIBERO-X, 49.7% overall on EBench, and 67.5% mean subtask success on a real dual-arm egg-cake task

Related research and updates

Synopsis

Vela is a vision-language-action foundation model pretrained in trajectory space that represents future motion with a fixed number of cubic B-spline control points plus a motion-dependent temporal span; pretrained on roughly 20,000 hours of public and roughly 20,000 hours of private robot data, it reaches 45.3% average success on LIBERO-X, 49.7% overall success and a 66 task-progress score on EBench, and 67.5% mean subtask success on a real wheeled dual-arm egg-cake cooking task.

Source-provided article image: Vela: Scaling Vision-Language-Action Models with Adaptive Action Curve Parametrization
Figure 1 ·

Figure 1: Overview of Vela. At inference, Vela jointly predicts a fixed set of spline control points and a motion-dependent horizon, defining a continuous action trajectory that can be sampled at arbitrary control rates. During training, motion-dependent horizon adaptation constructs supervision targets that adequately represent the demonstrated motion. The same control point budget thus covers longer spans for smooth motion and shorter spans when finer temporal resolution is needed.

arXiv

Interpretation

Representing future motion as a continuous trajectory rather than action chunks on a fixed temporal grid lets one output budget cover both long horizons and locally precise motion. Prior spline-based policies mainly target task-specific policies, single platforms, or downstream action tokenization and execution; this work makes trajectory-space action modeling the pretraining output space of a foundation model with a shared cross-embodiment action interface. On LIBERO-X the average success rate is 45.3% versus 39.3% for the point-wise baseline; on EBench the overall success rate is 49.7% with a progress score of 66, versus 41.4% and 54 for the baseline; ablations show post-hoc spline fitting alone (38.9% and 40.0%) does not recover the gains.

Motion-dependent horizon adaptation lets a fixed control-point budget allocate temporal resolution by motion complexity: longer spans for smooth motion, shorter spans for rapid local variation. Existing spline policies typically predict over a globally fixed temporal span, imposing one compromise between temporal coherence and local expressivity on all motions; here each demonstrated motion is assigned the longest span that still adequately represents it under a noise-aware criterion. With the same spline representation and control-point budget, the fixed-horizon variant averages 43.2% success versus 45.3% with adaptive horizons, improving at every one of the five perturbation levels.

Decoded-trajectory flow matching (DT-FM) weights control-point errors by their effect on the decoded trajectory, aligning supervision with trajectory geometry better than isotropic control-point supervision. CP-FM supervises the same flow residual with Euclidean distance in parameter space, whereas DT-FM measures it after spline decoding, equivalently replacing isotropic weighting with the decoder-induced quadratic form. CP-FM averages 44.0% success, an equal-weight CP-FM plus DT-FM combination 44.9%, and DT-FM alone 45.3%; the authors note the differences are modest, indicating effective trajectory-space learning does not hinge on a particular loss composition.

Trajectory-space modeling transfers to physical manipulation, remaining effective on multi-stage bimanual coordination, tool use, and sustained contact tasks. Most spline-policy studies stay in simulation or on a single platform; this work provides physical-robot evidence on a wheeled dual-arm robot for egg-cake cooking and potato shredding. Each of the four egg-cake subtasks is evaluated over 10 trials, with Vela at 67.5% mean subtask success versus 37.5% for the point-wise baseline and 40.0% for OpenWAM; on potato shredding Vela matches or exceeds both baselines at every stage.

Perspective

The results target imitation-learning-based robot manipulation with multi-embodiment demonstration data: pretraining uses roughly 20,000 hours of public data (60 datasets, 11 embodiment classes) plus roughly 20,000 hours of private data, at a scale of a 3B PaliGemma vision-language backbone with a 300M flow-matching action expert and no additional trainable parameters. The method suits tasks needing long-horizon coverage and contact-rich fine manipulation, such as tabletop precision assembly, mobile pick-and-place, and multi-stage bimanual cooking; at deployment the policy predicts control points and horizon directly from context without fitting or horizon search, so inference cost is comparable to the point-wise baseline. Teams wanting to reuse the representation can adopt its shared action interface and noise-calibrated target construction to convert existing action-chunk data into spline supervision targets.

Several points remain worth watching: first, the horizon-adaptation gain appears at all five perturbation levels but is modest in size (43.2% fixed versus 45.3% adaptive), and its stability across data scales and embodiments is not yet established; second, the three geometry objectives differ only slightly in average success (44.0%, 44.9%, 45.3%), so loss design reads as a refinement rather than a prerequisite; third, the real-world conclusions rest on two tasks with 10 trials per subtask, a limited statistical grain; fourth, the current policy adapts only the temporal horizon while keeping the control-point count fixed, and varying both would require variable-length prediction and matching training objectives, which the authors flag as an open direction; fifth, the model scale matches the point-wise baseline, so how trajectory-space modeling behaves with larger backbones and more data remains to be characterized.

Sources