Public articles linked to the same research event.
arXiv The work proposes UniIntervene++, which formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a semi-Markov decision process and learns their relative values online, combines periodic unassisted probing with coupled experience learning, and achieves 89.67% average success across five real-world manipulation tasks, at least 6 percentage points above the best baseline, while reducing the human intervention rate to 0.77%.
The work proposes UniIntervene++, which formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a semi-Markov decision process and learns their relative values online, combines periodic unassisted probing with coupled experience learning, and achieves 89.67% average success across five real-world manipulation tasks, at least 6 percentage points above the best baseline, while reducing the human intervention rate to 0.77%.
The work proposes UniIntervene++, which formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a semi-Markov decision process and learns their relative values online, combines periodic unassisted probing with coupled experience learning, and achieves 89.67% average success across five real-world manipulation tasks, at least 6 percentage points above the best baseline, while reducing the human intervention rate to 0.77%.
The work proposes UniIntervene++, which formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a semi-Markov decision process and learns their relative values online, combines periodic unassisted probing with coupled experience learning, and achieves 89.67% average success across five real-world manipulation tasks, at least 6 percentage points above the best baseline, while reducing the human intervention rate to 0.77%.