Skip to main content
Back to timeline
arXivSource publication:

EgoTools pairs 100 hours of first-person tool-use video with a 1,000-question benchmark, exposing video models' weak tool perception and causal reasoning

Synopsis

The work introduces EgoTools, a suite comprising EgoTools-Data, roughly 100 hours of tool-centric egocentric recordings across kitchen, classroom, research lab, repair workshop, craft, office, and household domains with synchronized audio, 361,332 hierarchical captions, 6,519 tool-centric narrations, and supplementary 3D information, plus EgoTools-Bench, a 1,000-question 8-way diagnostic benchmark of 900 human-crafted and 100 human-verified spatial questions across Affordance & Causality, Perception & Grounding, Procedural Dynamics, and Spatial Reasoning; evaluation shows Gemini-3.1-Pro at 66.9% overall but only 51.7% on Perception & Grounding, open-source instruct models between 38.9% and 50.

AI-generated editorial illustration: EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

Interpretation

EgoTools makes the tool the central unit of explanation in egocentric video understanding and supplies matching data and a diagnostic benchmark. Existing egocentric datasets and benchmarks are organized around activities, objects, events, or plans, and actions are often labeled at the level of "cook eggs" while the spatula and pan mediating the activity go unmentioned; EgoTools-Data instead records tool choice, substitution, target-object state change, and causal effects through tool-centric narrations, and EgoTools-Bench isolates tool-mediated reasoning for evaluation. The corpus contains 646 videos totaling 100.37 hours, with 361,332 hierarchical captions and 6,519 tool-centric narrations; the benchmark has 1,000 8-way questions (900 human-crafted plus 100 human-verified spatial), distributed as 363, 236, 222, and 179 across the four tracks, with an average temporal span of 4.19 minutes per question.

The benchmark exposes how current multimodal video models struggle to ground tool use in visual evidence, and this difficulty is not an isolated weakness of one architecture. Models that perform well on existing egocentric benchmarks such as EgoSchema and EgoThink drop noticeably on EgoTools-Bench, indicating that broad first-person activity understanding does not automatically yield tool-use reasoning. Chance on the 8-way format is 12.5%; open-source instruct models range from 38.9% to 50.0% overall, with Qwen3-VL-8B-Instruct strongest at 50.0%; among open-source thinking models GLM4.1V-Thinking leads at 50.7%; Gemini-3.1-Pro reaches 66.9% overall but only 51.7% on Perception & Grounding, 18-23 points below its AC, PD, and SR scores; human experts reach 83.2% overall.

EgoTools-Data is not merely an intermediate source for benchmark construction but actionable supervision that partially closes the tool-centric reasoning gap. With training and benchmark pools separated at the source-video level, single-stage full supervised fine-tuning on 184,679 instruction examples derived from the non-benchmark portion still improves accuracy on held-out tool-use episodes, indicating generalization rather than memorization of benchmark clips. Qwen3-VL-8B-Instruct improves from 50.0% to 60.9% (+10.9 points), with Affordance & Causality +9.3, Perception & Grounding +13.8, and Procedural Dynamics +11.0, while Spatial Reasoning drops 11.0 points; the same model reaches 67.6% on EgoSchema, 1.4 points below the 69.0% backbone.

Input ablations separate the diagnostic roles of the four tracks: perception grounding depends on visual evidence, procedural dynamics on temporal order, and spatial reasoning on multi-frame coverage. By contrasting ordered 64 frames, shuffled 64 frames, a single middle frame, and text-only input, the work attributes the needs for vision, order, and multi-frame coverage to different tracks rather than claiming generically that models need video. Ordered 64 frames reaches 51.7% overall versus 42.7% for text-only (a 9.0-point gap, with text-only still above the 12.5% chance level); the middle frame alone drops to 46.9%, with an 11.9-point gap on Spatial Reasoning; shuffling frames drops to 49.3%, with Procedural Dynamics most order-sensitive (3.2-point gap) and Spatial Reasoning only 1.1 points apart.

Perspective

The work targets researchers and system developers who need to understand tool-mediated action, and applies to question answering and training over tool selection, substitution, manipulation order, and target-object state change in egocentric video. EgoTools-Data spans kitchen, classroom, research lab, repair workshop, craft, office, and household domains and can be used to build instruction-tuning corpora; EgoTools-Bench is for diagnostic evaluation, with its four tracks covering perception grounding, procedural dynamics, spatial relations, and affordance causality. The authors state the benchmark is dominated by kitchen questions and should not be treated as domain-balanced, so domain-level conclusions should be used cautiously; public release is scoped to research use and includes rectified egocentric videos, hierarchical captions, corrected English narrations, all 1,000 questions with evaluation annotations, clip metadata, and source-video-disjoint split information, while raw fisheye videos are not released and narration audio, selected grounding annotations, and 3D trajectory assets are released incrementally subject to authorization and privacy review.

The benchmark's domain distribution is strongly long-tailed and dominated by kitchen questions, so domain-level comparisons would be unstable, which is why the authors focus primary analysis on reasoning capabilities rather than domain performance. Spatial Reasoning drops 11.0 points after fine-tuning, indicating that spatial reasoning remains an open problem under the current training recipe; on EgoPlan-Bench the model reaches 35.2%, 7.1 points below the 42.3% backbone, and the authors note the training subsets mirror those benchmarks' answer formats, so format coverage cannot be separated from other fine-tuning effects. Text-only input still reaches 42.7%, above the 12.5% chance level, suggesting some questions retain exploitable commonsense or answer-option cues. Model variants also differ in training and sometimes input budgets, so the effects of scale, audio, or deliberation are not isolated; no formal institutional ethics review was sought for data collection, with informed consent, anonymization, and release review described as the safeguards instead.

Sources