Skip to main content
Back to timeline
arXivSource publication:

UniIntervene++ raises average success to 89.67% across five real robot manipulation tasks while cutting human intervention to 0.77%

Related research and updates

Synopsis

The work proposes UniIntervene++, which formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a semi-Markov decision process and learns their relative values online, combines periodic unassisted probing with coupled experience learning, and achieves 89.67% average success across five real-world manipulation tasks, at least 6 percentage points above the best baseline, while reducing the human intervention rate to 0.77%.

Source-provided article image: UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning
Fig. 1 ·

Fig. 1: Competence-adaptive control allocation in UniIntervene++ . Illustrative Ring Transfer rollouts show a shift from early failures with substantial use of Trajectory Correction and CodePolicy to later autonomous success with predominantly RL policy control. Colored bars indicate Option occupancy within each rollout (teal: RL policy; orange: Trajectory Correction; purple: CodePolicy). Assisted experience improves the RL policy, while Adaptive RL Probe reserves unassisted windows to assess its evolving competence and guide subsequent allocation.

arXiv

Interpretation

It formulates online intervention as a unified control-allocation problem, placing the evolving RL policy, demonstration-guided trajectory correction, and a task-structured CodePolicy in a shared decision space and learning their relative values online with an availability-masked SMDP Double-DQN scheduler. AutoSERL relies on trajectory guidance and predefined failure regions from a single demonstration, and UniIntervene triggers an offline-learned recovery policy on value degradation; this work moves the intervention learner into the online loop so that entry, corrective behavior, and exit all follow the same evolving value comparison. The method specifies the Option set, scheduler state, availability mask, scheduler reward, and SMDP target, and states that the underlying HIL-SERL off-policy objective and update rule are retained unchanged.

It introduces competence-adaptive intervention: Adaptive RL Probe periodically reserves a contiguous unassisted window, forces the RL Option at each boundary within it, and judges probe passes by projected arc-length progress along the demonstration reference path together with a maximum path deviation, adjusting the unassisted window accordingly. The probe adds no scheduler action, reward, or value function; its progress, deviation, and pass signals only update the curriculum, are not added to the scheduler reward, and a probe pass is not counted as an independent task success, so allocation follows measured execution evidence rather than assisted-episode success. The pass condition is given explicitly, and the text describes how task completion, human takeover, safety interruption, excessive path deviation, or stalled progress each end a probe.

It introduces coupled experience learning: all robot-control transitions enter the online replay buffer, while transitions generated by automated assistance or human override are additionally routed to the demonstration buffer, from which the unchanged HIL-SERL learner draws mixed off-policy batches. Assisted experience contributes directly to RL policy learning without an additional behavior-cloning or policy-distillation objective; the scheduler and RL policy share no gradients and are coupled instead through a data loop across robot-control and Option timescales. The routing rule is stated explicitly, and the scheduler and RL learner are described as retaining separate networks, optimizers, and replay buffers.

Across Ring Transfer, Tube Insertion, USB Insertion, Wipe Whiteboard, and Towel Folding, UniIntervene++ reaches 89.67% average success with a 0.77% human intervention rate; on Ring Transfer the RL policy's control share rises from 35.78% to 90.51% as its independently evaluated success rises from zero to 88.33%. Against the best baseline average of 83.67% (HIL-SERL and UniIntervene), success is 6 percentage points higher, and relative to AutoSERL, the baseline with the lowest human intervention rate, human-controlled steps fall by at least 94.6% relative. Success rate is measured over three evaluation runs of 20 episodes per task after training; ablations compare five configurations on Ring Transfer and Towel Folding, where the full method reaches 88.33% and 85%, the fixed-prior stochastic scheduler 75% and 80%, removing CodePolicy 0 on both, and removing Trajectory Correction 50% and 47%.

Perspective

The results apply to a real manipulation setting with a single UR7e arm and a Robotiq gripper, where human corrective actions and safety overrides are issued through a 3Dconnexion SpaceMouse, and cover precision manipulation, contact-rich insertion and interaction, and deformable-object manipulation. Reported performance is reached within a range of robot-interaction steps corresponding to a range of minutes of effective training time, with USB Insertion the shortest run and Ring Transfer the longest. Human takeover remains an external safety backstop, so the framework suits training pipelines that still permit human safety intervention; for teams that want routine correction and recovery handled by automated assistance while reserving human involvement for early unsafe states, the scheduler design is directly reusable.

The ablations are system-level: removing an Option changes both the assistance executed and the experience routed to the task policy, so the contribution of each assistance source cannot be fully separated, and all scheduler-enabled variants retain Adaptive RL Probe, whose independent contribution is not isolated. Success rates rest on three evaluation runs of 20 episodes per task, so the stability of task-level differences would benefit from further repetition. On Wipe Whiteboard, UniIntervene++ reaches 86.67%, below UniIntervene's 90.00%, suggesting that the relative advantage of adaptive allocation over fixed offline recovery may vary with the task, particularly for sustained tool-surface interaction. In addition, the probe pass condition depends on arc-length progress and deviation thresholds relative to the demonstration reference path, with a tighter local-deviation condition when the reference segment is nearly stationary, and the sensitivity of results to those thresholds is not explored in the text.

Sources