Skip to main content
Back to timeline
arXivSource publication:

CARE turns robot execution failures into training data, lifting average task success by 14.5 points in simulation and 15.9 points in the real world

Synopsis

The work proposes CARE, which collects failed rollouts, models stage-conditioned post-failure geometric deviation distributions, synthesizes representative failure states and corrective demonstrations from them, and at inference combines stage-wise planning with physically grounded 3D point-cloud monitoring to trigger atomic adjustments or re-operations; across RoboTwin 2.0, RoboFactory, the newly introduced FSR-Bench, and real-world dual-arm tasks it reports average task-success gains of 14.5 points in simulation, 15.9 points in the real world, and a 7.5-point gain in recovery success on FSR-Bench.

AI-generated editorial illustration: CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

Interpretation

CARE treats execution failures themselves as structured supervision: it first executes nominal atomic plans without corrective supervision, collects failures from 100 rollouts where an atomic stage does not satisfy its termination condition, characterizes stage-conditioned post-failure deviations as relative translational and yaw offsets, samples from that distribution, perturbs the stage-relevant relative pose, rolls forward under physical dynamics to obtain new physical failure states, and collects short corrective demonstrations via an IK/motion-planning oracle in simulation or human teleoperation in the real world. Prior approaches largely collect corrective demonstrations from manually designed or random perturbations, or treat failures as exceptions to be rolled back; CARE explicitly models the stage-conditioned post-failure deviations induced by the policy's own execution and uses them to guide corrective data synthesis. The paper reports failure modeling from 100 preliminary rollouts and selects among Gaussian, Beta, Gamma, Weibull, Log-Normal, and Uniform families per scalar deviation variable using goodness-of-fit statistics with AIC; ablations show 50 rollouts are consistently weaker while 100 and 150 differ little, with JS divergences to the full-data distribution at 100 rollouts of 0.0224, 0.0200, 0.0374, and 0.0245.

CARE decomposes long-horizon manipulation into atomic stages, uses GPT-4.1 at temperature 0 queried once at task initialization to produce stage plans containing a nominal atomic instruction, geometric cues, a stage-termination predicate, and intra-execution adjustment and post-execution re-operation instructions, and at execution time uses a point-cloud predicate checker as a 3D monitor triggered at stage-critical transitions such as grasp closure or object release, extracting gripper, object, and target masks with SAM 3 and depth with Depth Anything 3 to compute a compact geometric state and select continue, adjust, or re-operate. Unlike methods relying on VLM judgment, world models, or rollback, CARE uses skill-level geometric predicate templates rather than task-specific learned classifiers, and the same VLA executor handles nominal, adjustment, and re-operation instructions without switching to a separate recovery policy. The monitoring ablation compares a GPT-4.1-based VLM monitor with the 3D monitor on the same triggered observations across three simulated and two real-world tasks, raising average accuracy from 71.4% to 93.8% (+22.4 points) and reducing end-to-end monitor latency from 3.40 s to 0.24 s; across a broader task list the monitor averages 93.2% accuracy.

The paper introduces FSR-Bench, which evaluates recovery from intermediate failure states rather than clean initial states, containing 36 recovery scenarios across five tasks split into 21 Easy local failures typically admitting a single corrective operation and 15 Hard structural failures requiring multi-step correction or coordinated multi-arm recovery; its recovery training data are uniformly sampled perturbations within predefined ranges, while evaluation initializations come from stage-conditioned empirical failure distributions estimated from nominal-policy rollouts and rolled forward under physical dynamics. Existing manipulation benchmarks largely measure task execution from predefined initial states and offer limited support for evaluating recovery from intermediate failures; FSR-Bench reports recovery success independently from nominal task execution and deliberately decouples the training distribution from the test distribution. Recovery success rate is computed over 100 trials per task-regime pair, and all baselines and CARE variants use the same uniformly sampled corrective training data and the same pre-generated failure initializations, so the comparison isolates atomic corrective execution.

Across multiple VLA backbones, simulation benchmarks, and real-world dual-arm tasks CARE yields consistent gains: on RoboTwin 2.0, π0 rises from 39.6% to 60.7% (+21.1 points) and π0-FAST from 17.4% to 33.1% (+15.7 points); on RoboFactory, average absolute gains are +9.3 points for π0-FAST and +12.0 points for π0; on FSR-Bench, π0's average recovery success rises from 38.8% to 57.2% in the Easy regime and from 15.8% to 21.2% in the Hard regime; on the real SO-101 dual-arm platform, π0 rises from 30.8% to 50.0% and π0-FAST from 18.8% to 31.3%, exceeding the reproduced FailSafe recovery baseline by 13.7 and 7.3 points respectively. These results extend the claim that execution failures provide effective corrective supervision from a single backbone and benchmark to continuous and tokenized action-modeling backbones, multi-arm collaboration, a recovery-specific benchmark, and real-world deployment. Simulation and real-world evaluations use matched data budgets (for example, 150 nominal expert demonstrations plus 50 additional trajectories per error type in standard simulation tasks, and 50 nominal demonstrations plus 20 additional trajectories per error type per real-world task) with 100 real-world trials per task; component ablations show the largest gains when data and execution are combined, and experience-guided sampling outperforms uniform random sampling by 12.2 points on π0 and 8.8 points on π0-FAST on average.

Perspective

The work targets long-horizon multi-arm VLA manipulation where tasks have well-defined atomic stages and stage-critical geometric events such as grasp closure or object release, and it is validated on RoboTwin 2.0, RoboFactory, FSR-Bench, and a dual-arm LeRobot SO-101 platform. It lets follow-up work reuse the failure-distribution-modeling plus atomic-corrective-execution pipeline, the FSR-Bench recovery evaluation protocol, and skill-level geometric predicate templates; for researchers and engineering teams wanting to improve recovery without replacing the backbone model, this path turns recovery capability into reusable components at the data and execution-mechanism level. The paper also states that future work will integrate reinforcement learning into CARE to further enhance recovery.

The 3D monitor depends on SAM 3 segmentation and Depth Anything 3 depth estimation, so its behavior on new objects and scenes is worth watching; geometric tolerances are calibrated from rollout statistics, so transferring to a new atomic skill requires defining its predicate template and recalibrating numerical values. Hard structural-failure recovery on FSR-Bench remains low overall (for example, 21.2% for π0), indicating that multi-step correction and coordinated multi-arm recovery remain open problems. In addition, the failure distribution is estimated with 100 preliminary rollouts as a practical trade-off, so whether that budget remains sufficient across tasks and platforms, and how reinforcement learning would divide labor with existing corrective supervision once integrated into CARE, are questions to track.

Sources