HuRo turns five human-video sources into about 142 million robotized frames, lifting real-robot OOD completion from 34.9% to 72.2% in ALLEX VLA pretraining
Synopsis
HuRo introduces a robotization pipeline that converts heterogeneous egocentric human videos into robot-aligned observations and action trajectories, builds the HuRo dataset of about 142 million frames from five sources (EgoDex, EgoVerse, Ego4D, Ego10K, and EPIC-Kitchens), and uses it to pretrain VLA policies on the ALLEX bimanual dexterous robot: as robotized data increases, average completion across four real-world manipulation tasks rises from 51.5% without pretraining to 80.3% with the full dataset, with ID completion from 68.1% to 88.4% and OOD completion from 34.9% to 72.2%.
Interpretation
The paper builds a joint observation-action robotization pipeline: human-video annotation (camera intrinsics, hand detection and MANO hand pose, masked DROID-SLAM camera trajectory with MoGe-2 metric scale and GeoCalib gravity alignment, chunking and Qwen3.5 language instructions), action conversion (two-stage PyRoKi retargeting of human hand motion into robot joint trajectories, with action annotations derived from the state trajectory), and visual conversion (SAM2 arm segmentation, ProPainter inpainting, and Isaac Sim rendering of the robot overlay). Prior human-to-robot work has largely been demonstrated in task-matched settings or has addressed observation and action alignment separately; this pipeline uses available annotations and estimates missing signals to unify heterogeneous sources with different annotation levels into a common robot observation-action format. The method is described stage by stage in the main text and appendix, with components and two-stage objectives; Appendix F reports pipeline cost and quality diagnostics, including median rotation error and 4.47 mm median position ATE for camera trajectory against EgoDex annotations, 20.4 mm median root-relative 21-keypoint hand-pose error, and 21.2 mm median fingertip residual after retargeting.
The resulting HuRo dataset covers five egocentric human-video sources, by frame count roughly 55.5% EgoDex, 26.5% EgoVerse, 10.4% Ego4D, 6.0% Ego10K, and 1.7% EPIC-Kitchens, totaling about 142 million frames and about 1,316.8 hours. The paper states this is over an order of magnitude larger than data used in prior robotized-video pretraining methods, and it is mixed-source rather than single-source. Appendix D gives per-source frame counts, hours, and frame ratios; Appendix D.2 measures visual coverage with DINOv3 features against an OpenImages reference set and instruction coverage by unique verbs, objects, and verb-object pairs, showing coverage rising with mixed-source subset size and mixed 50% covering more than EgoDex-only at similar sampling.
On four real-world ALLEX manipulation tasks, downstream performance improves as the amount of robotized human-video pretraining data increases: average completion rises from 51.5% without pretraining to 80.3% with the full dataset, ID from 68.1% to 88.4%, and OOD from 34.9% to 72.2%; the full-scale model also outperforms the reported reference models under both ID and OOD evaluation. Prior work has used robotized videos mainly for visual representation or auxiliary prediction objectives; this paper treats robotized data as a scalable source of joint observation-action supervision and provides data-scaling evidence. The four tasks use 43, 40, 16, and 20 demonstrations, with ID/OOD rollouts of 12/12, 12/24, 12/12, and no ID/10; Apple Pick-and-Place uses binary success while the others use stage-wise partial-completion scores. Pretrained variants fix optimization steps and global batch size while varying only data amount.
Ablations show separate contributions from visual robotization and retargeted action supervision: the no-overlay variant matches full HuRo on ID (89.4% vs. 88.4%) but is substantially lower on OOD (55.7% vs. 72.2%); on Diverse Pick-and-Place, visual-only transfer gives only modest gains over no pretraining, whereas end-to-end pretraining with retargeted actions reaches higher completion under both ID and OOD. The paper separates visual conversion from action supervision, showing they are not substitutes and that action supervision adds benefits beyond visual-only transfer. The no-overlay variant uses the same pretraining data and retargeted action supervision but retains original human-video observations; the action-supervision comparison includes No PT, PT (Visual Only, with the action head reinitialized), and PT (Visual + Action), evaluated over 36 ID and 18 OOD trials.
Perspective
The results are aimed at robot-learning researchers and engineering teams who want to expand VLA pretraining data with human video, in settings such as tabletop manipulation on a bimanual dexterous robot like ALLEX, and on OpenArm as a transfer target in the appendix. The paper states that robotized observation fidelity is bounded by reconstruction and visual conversion quality, and that the current robot overlay does not explicitly model occlusion between the rendered robot and scene geometry; the dataset provides visual and kinematic action supervision but does not capture force or tactile signals; kinematic retargeting does not model self-collision or physical contact, and in a five-source audit only 55.2% of trajectories had no detected non-grasp self-contact in sampled frames, so the trajectories are positioned as pretraining supervision rather than executable robot demonstrations.
Several numbers in the main text appear as placeholders in the loaded text (for example dataset scale, completion start and end values, and some hyperparameters), so the concrete figures in this summary are drawn mainly from the appendix tables; readers citing main-text values should check the original paper. The paper states that how different levels of robotization fidelity affect downstream policy learning remains unexplored, so the fidelity-performance relationship is an open question. The absence of self-collision, contact force, and tactile signals means behavior on contact-rich tasks still needs further observation. Benefits from multi-embodiment data are only preliminarily compared in the appendix, and how embodiment diversity, morphology, and action interfaces affect the utility of robotized data is listed as future work.
