Skip to main content
Back to timeline
arXivSource publication:

VLA-Precision reaches 98.3% mean success across nine precision chemistry tasks with 45.8 minutes of online training per task

Synopsis

The work presents VLA-Precision, a framework for real-world online reinforcement learning of large vision-language-action (VLA) models, combining the Asymmetric Co-Bootstrapping (ACoB) algorithm with the ACoB-Stream training architecture, and reports 98.3% mean success across nine high-precision chemistry tasks in four categories and four robot platforms, with 45.8 minutes of online training per task, 27.6-second episodes, and up to 10.9x improvements in throughput and computational efficiency.

AI-generated editorial illustration: VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

Interpretation

ACoB couples rapid behavioral learning with progressive value calibration across timescales: early intervention-guided behavior cloning rapidly absorbs successful actions and human corrections, while global return propagation and local preference ranking continually calibrate value estimates as autonomous experience accumulates, yielding relative action advantages for reference-regularized policy improvement. Prior real-world VLA RL largely uses behavior cloning to constrain the policy distribution for stable exploration, which offers limited help to value estimation itself; ACoB instead uses behavior cloning to accelerate experience quality and adds state-matched local preference ranking to correct the value of overwritten proposals. Ablations on four representative tasks provide the comparison: full ACoB reaches 96.25% final autonomous success with a 0.03% intervention rate; removing critic preference drops it to 26.25%, removing actor BC to 20.00%, and replacing relative advantage with direct Q maximization to 8.75%, with intervention rates rising to 10.57%, 9.24%, and 24.93% respectively.

ACoB-Stream adopts invariant-state decoupling and on-demand streaming as design principles, managing the lifecycles of experience-context state and policy state across four stages: experience-context formation, persistence, access, and policy-state synchronization. Earlier asynchronous actor-learner systems for compact policies rely on replay-intensive optimization whose cost grows with large VLAs; ACoB-Stream instead reuses frozen-prefix KV contexts across updates, uses context deduplication with disk-backed sliding-window sampling and objective-aligned context retrieval, and publishes only the trainable action-expert state rather than synchronizing the full model. In the system comparison, Disk ACoB-Stream completes 15,666 CTA cycles in 150 minutes at 1.7407 CTA/s and 0.574 s per cycle; versus No KV Cache it delivers 10.95x throughput and 90.9% lower mean latency; versus Naive Context Access 1.80x throughput and 44.4% lower mean latency; versus Dynamic CPU RAM 1.09x throughput and 8.0% lower latency.

Across nine high-precision chemistry manipulation tasks, VLA-Precision achieves 98.3% mean success with 45.8 minutes of online training per task and 27.6-second average episodes, succeeding in 177 of 180 held-out trials, reaching 100% on seven tasks and at least 90% on every task. On the same task suite, HIL-SERL averages 2.8% success, ConRFT 7.8%, and Robo-Dopamine 10.0%, while the two demonstration-only large VLA baselines reach 59.4% and 67.8%; relative to the strongest real-world RL baseline, Robo-Dopamine, VLA-Precision improves average success by 88.3% and execution speed by 61.2%. Results come from real-world experiments on four robot platforms (two UR5e single-arm configurations, a UR5e dual-arm setup, and a Franka Research 3) across contact-rich, contact-light, contact-free, and bimanual coordination task categories, with object poses and initial arm poses deliberately randomized in every task.

Offline comparisons show online RL improves task policies beyond demonstration-only fine-tuning: VLA-Precision succeeds in 99 of 100 trials across five representative tasks, versus 61 and 50 for the two demonstration-only baselines, with 26.65 s versus 29.32 s and 32.05 s average successful-trial time. VLA-Precision uses fewer demonstrations (60-120 versus 120-200) and fewer Stage-I steps (5,000-15,000 versus 25,000-30,000), indicating the gain comes from Stage-II online RL rather than more supervised data. This comparison is an offline test of 20 trials per method per task; on the bimanual tube-brushing task the two baselines fall to 10% and 5% while VLA-Precision reaches 95%.

Perspective

The work targets pretrained large VLAs that already have task-level competence, performing online post-training on real robots in precision manipulation settings with tight positional and angular tolerances, such as tube-rack loading, pipetting, and stopper insertion in chemistry laboratories. It offers a task-specific improvement path: full-parameter imitation learning on demonstrations to obtain a task prior, then freezing the backbone and training only LoRA parameters in the action expert. For teams aiming to reduce human intervention, shorten per-task online training time, and sustain update throughput under limited compute and memory, ACoB and ACoB-Stream provide reusable algorithmic and system design references. The authors also list long-horizon bimanual real-world RL and multi-task joint optimization as future directions, and note that step-wise delta action representations accumulate prediction error as the execution horizon grows, motivating chunk-wise deltas for long-horizon bimanual manipulation.

All results come from the authors' own high-precision chemistry task suite and four robot platforms; although the tasks are numerous and varied, they remain a specific setting, and the compared baselines are reproduced or cited from prior work whose tuning budgets may not be identical. Ablations are run on four representative tasks, and the system comparison is measured on specific hardware within a 150-minute window, so throughput and latency numbers will vary with hardware, data scale, and implementation details. Only one long-horizon bimanual task is included, and the authors themselves note that error accumulation over longer sequences is not fully explored; multi-task joint optimization, task-conditioned value learning, and catastrophic-forgetting control remain open questions. This summary is based on the abstract and visible body text, and per-task curve details in the figures are not enumerated, so readers needing exact numbers for reproduction should consult the original figures and tables.

Sources