Skip to main content
Back to timeline
arXivSource publication:

InfiniHand estimates hands and camera trajectory together from egocentric video with one streaming feed-forward network, cutting ARCTIC PA-p to 7.72 mm at 11.19 FPS

Synopsis

InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.

AI-generated editorial illustration: InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video

Interpretation

The work unifies hand localization, MANO parameter prediction, and camera trajectory estimation, previously handled by cascaded independent modules, into a single streaming feed-forward architecture that outputs world-space hand motion directly from uncalibrated egocentric video. The paper argues that existing approaches such as HaWoR, which combines hand detection and tracking, a camera-space hand reconstruction network, and DROID-SLAM with Metric3D, rely on cascading independent estimators and SLAM, causing error accumulation, complex pipelines, and computational overhead; InfiniHand instead explicitly couples global camera motion with local hand articulation in a shared spatiotemporal representation. The abstract and introduction describe the architectural unification, and the method details mask-based hand localization, dual-stream feature fusion, MANO regression, and joint camera decoding; the ablation shows that omitting Stage II raises W-MPJPE from 59.21 to 185.76 mm, supporting the role of joint modeling for world-space estimation.

For camera-space reconstruction, InfiniHand achieves the lowest PA-p across four datasets, reducing ARCTIC from ViDiHand's 9.82 mm to 7.72 mm (21.4%) and EgoDex from 17.22 mm to 7.29 mm (57.7%). Relative to methods such as ViDiHand that adapt video diffusion priors, the gain is attributed to fusing hand-centered appearance features with geometric features and solving translation via differentiable least-squares projection. Table 1 covers ARCTIC, HOT3D, EgoDex, and HOI4D against InterWild, HaMeR, Hamba, WildHands, OmniHands, WiLoR, Dyn-HaMR, HaWoR, and ViDiHand; the ablation shows that removing WiLoR features raises MP-p from 17.09 to 33.31 mm and replacing the learned mask head with HaWoR masks raises PA-p from 7.72 to 25.50 mm.

For world-space reconstruction, InfiniHand achieves the lowest W-MPJPE on all three datasets, reducing ARCTIC from 65.86 to 59.21 mm (10.1%) versus WiLoR-SLAM and EgoDex from 78.41 to 28.77 mm (63.3%) versus Dyn-HaMR. The paper attributes this to a sparse bundle adjustment backend beyond streaming trajectory memory, which uses a binary keyframe pool to revisit past observations and curb long-sequence drift. Table 2 compares against HaWoR, Dyn-HaMR, and WiLoR-SLAM; the ablation shows that disabling sparse BA raises W-MPJPE from 59.21 to 80.48 mm while camera-space metrics stay unchanged, indicating its effect is concentrated in global reconstruction.

On efficiency, InfiniHand reaches 11.19 FPS, above WiLoR-SLAM at 8.53 FPS, HaWoR at 5.48 FPS, and Dyn-HaMR at 0.81 FPS, more than twice the throughput of HaWoR. The paper explains this comes from reusing streaming features and overlapping execution across temporal windows: two primary workers process the upcoming 16-frame window and decode the prior window, while a third worker runs sparse BA in the background when needed. The efficiency comparison is reported under standard timing protocols with an implementation description of the pipelined execution; the paper also notes that background BA still incurs additional computation.

Perspective

The result targets research and engineering settings that need world-space hand motion from egocentric video, for example generating 3D hand demonstrations for embodied learning, human-to-robot motion transfer, or manipulation policy learning. The evaluation covers camera-space benchmarks on ARCTIC, HOT3D, EgoDex, and HOI4D, world-space benchmarks on ARCTIC, HOT3D, and EgoDex, and qualitative in-the-wild assessment on Ego4D; training uses roughly 5,000 hours aggregated from nine public egocentric datasets, uniformly sampled at 10 FPS. Methodologically it builds on the streaming 3D foundation model LingBot-Map and keeps a lightweight sparse bundle adjustment as post-optimization, so it fits deployments that can accept streaming inference plus periodic background refinement.

The limitations the paper itself states are worth watching as follow-up questions: the underlying LingBot-Map lacks inherent metric scale and requires an auxiliary post-processing alignment model whose errors can propagate to hand positions and global trajectories; high-quality egocentric data with large-amplitude two-hand interactions and camera dynamics remains scarce, limiting generalization to rapid viewpoint changes, severe occlusions, and intermittent hand visibility; and monocular pipelines often inherit time-varying scale drift from monocular annotations, affecting long-sequence trajectory accuracy and temporal smoothness. In addition, this reading covers the paper text and appendices, with figures described in words rather than viewed as images, so qualitative comparisons such as finger articulation under occlusion or two-hand orientation agreement with reference trajectories can only be understood from the text; the visual details still require the original Figures 3 to 6.

Sources