Public articles linked to the same research event.
arXiv Vela is a vision-language-action foundation model pretrained in trajectory space that represents future motion with a fixed number of cubic B-spline control points plus a motion-dependent temporal span; pretrained on roughly 20,000 hours of public and roughly 20,000 hours of private robot data, it reaches 45.3% average success on LIBERO-X, 49.7% overall success and a 66 task-progress score on EBench, and 67.5% mean subtask success on a real wheeled dual-arm egg-cake cooking task.
Vela is a vision-language-action foundation model pretrained in trajectory space that represents future motion with a fixed number of cubic B-spline control points plus a motion-dependent temporal span; pretrained on roughly 20,000 hours of public and roughly 20,000 hours of private robot data, it reaches 45.3% average success on LIBERO-X, 49.7% overall success and a 66 task-progress score on EBench, and 67.5% mean subtask success on a real wheeled dual-arm egg-cake cooking task.
Vela is a vision-language-action foundation model pretrained in trajectory space that represents future motion with a fixed number of cubic B-spline control points plus a motion-dependent temporal span; pretrained on roughly 20,000 hours of public and roughly 20,000 hours of private robot data, it reaches 45.3% average success on LIBERO-X, 49.7% overall success and a 66 task-progress score on EBench, and 67.5% mean subtask success on a real wheeled dual-arm egg-cake cooking task.
Vela is a vision-language-action foundation model pretrained in trajectory space that represents future motion with a fixed number of cubic B-spline control points plus a motion-dependent temporal span; pretrained on roughly 20,000 hours of public and roughly 20,000 hours of private robot data, it reaches 45.3% average success on LIBERO-X, 49.7% overall success and a 66 task-progress score on EBench, and 67.5% mean subtask success on a real wheeled dual-arm egg-cake cooking task.