CAPABLE lifts a frozen VLA from 24.8% to 59.3% success on an actuator excluded from fault training
Synopsis
CAPABLE is a capability-aware adaptation framework for frozen vision-language-action policies: a temporal encoder shared across joints infers from command-response history and live Jacobian grounding how much commanded motion each joint still realizes, a self-supervised physical-prediction objective ties that representation to behavior, and a bounded residual reinforcement-learning policy corrects the VLA arm action without fault labels or faulty-joint identifiers; across 28 LIBERO tasks it raises success on an actuator excluded from fault training from 24.8% to 59.3%, 17.4 points above a parameter-matched global-history baseline, while preserving healthy performance.
Interpretation
The work reformulates actuator faults as remaining capability, meaning how much of a commanded motion each joint actually realizes and how that motion contributes to end-effector behavior in the current configuration, and uses it to correct a frozen generalist policy. Unlike fault-recovery methods that need task-specific retraining, fault labels, explicit diagnosis, faulty-joint supervision, or privileged embodiment information, CAPABLE uses only command-response history, controller signals, and known kinematics; no fault label, faulty-joint identifier, or fault-specific demonstration enters any network during training or deployment. Across 28 LIBERO tasks, success on an actuator excluded from fault training rises from 24.8% to 59.3%, 17.4 points above a parameter-matched global-history baseline, while healthy success remains comparable to zero correction.
The work introduces a per-actuator factorized architecture: one temporal encoder is shared across all seven joints, each joint's live Jacobian column grounds the estimate in the current configuration, and cross-joint attention composes the joints into a representation of what motions the arm can still produce. Prior weight-sharing work shares modules across morphologies to generalize over bodies; here a module is shared across actuators of one body to generalize over which actuator has failed, with the live Jacobian turning the shared estimate into a configuration-specific correction. Ablations show that removing history leaves a memoryless residual at 39.4%, below the unfactorized baseline; unsharing the temporal encoder costs 13.5 points, removing the Jacobian query 11.5, cross-joint attention 13.2, the self-supervised capability loss 15.6, and conditioning 9.7; the unshared variant performs worse with 50% more parameters.
The work separates factorization from capacity using a parameter-matched global-history baseline and tests whether transfer is specific to one actuator through leave-one-actuator-out experiments across six joints. Prior work evaluated held-out joints but did not establish systematic transfer for a frozen multi-task VLA without fault labels, affected-joint supervision, or healthy reference trajectories. CAPABLE outperforms the global-history baseline on all six held-out actuators, and leads on all four LIBERO suites and 24 of 28 task cells, with margins ranging from a few points upward.
The work characterizes where capability transfer holds and where it does not: it leads on persistent lock, viscous damping, Coulomb friction, and late-onset lock, but not on partial actuator effectiveness or restricted joint range. This indicates that what transfers is the trained signature of motion suppressed relative to command, whereas graded and state-dependent impairments differ from that signature in ways the current formulation does not capture. Cross-fault evaluation is zero-shot on five unseen fault families at three severity levels, with the same 300 held-out episodes per task-condition cell; on hardware, a physical Franka Panda under software-enforced joint locks gives CAPABLE 80.0% (24/30) overall versus 43.3% (13/30) for the global-history baseline, with the pooled difference significant under a two-sided Fisher exact test.
Perspective
The work targets an already-trained frozen VLA on a redundant manipulator: under faults that suppress motion relative to command, such as persistent lock, viscous damping, Coulomb friction, and late-onset lock, capability inference transfers zero-shot to a joint excluded from fault training. The intended setting is continuing a language-conditioned manipulation task after a fault instead of halting, with residual training in simulation and zero-shot transfer to hardware. For a reader, this means not retraining a task policy per fault and not supplying fault labels or faulty-joint identifiers, only available command-response history and known kinematics; hardware validation runs on a physical Franka Panda under software-enforced joint locks, with a digital twin matching robot/table geometry, camera pose, objects, controller, and control rate.
The cross-fault results show that partial actuator effectiveness and restricted joint range are not covered by the current formulation, since both are graded or state-dependent and differ from the trained persistent-lock signature; hardware evaluation covers only three task-joint conditions with 10 trials each, individual conditions are not significant at that sample size, and only the pooled difference is significant; the leave-one-actuator-out study uses eight rather than all 28 tasks; training uses persistent locks only. A careful reader would watch how extending capability inference to graded and state-dependent impairments changes the transfer boundary, and whether larger hardware trials and more task subsets preserve the same factorization advantage.
