Skip to main content
Back to timeline
arXivSource publication:

GEB links visually grounded observations into entity biographies, lifting EgoLifeQA accuracy to 72.0%, 4.4 points above the strongest published memory framework

Synopsis

The work introduces Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies and retrieves a biography alongside episodic evidence at question time; across four benchmarks including day-long and week-long recordings it improves both multiple-choice and open-ended question answering over prior memory frameworks, reaching 72.0% accuracy on EgoLifeQA, 4.4 percentage points above the best published result, with ablations showing that grounded identity association and biography reading each contribute and that additional descriptions alone do not fully recover the gains.

AI-generated editorial illustration: Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Interpretation

GEB organizes observations of the same physical instance into temporally ordered biographies, where each encounter preserves the visible subject, its event context, and supporting evidence, and entity nodes act as bridges across episodes. Prior memory frameworks rely on temporal descriptions and text-derived entities, where one description can cover different objects and one object can be split across names as its state changes; GEB instead associates observations across clips using visual and contextual evidence and vetoes a match when shared frames show visual separation. The paper gives a formal representation (entity, observation, episode, and source-clip nodes with three relation types) and builds memory over the EgoLife week: 6,266 thirty-second clips, 308,244 tracked observations, 77,716 entities after association, of which 27,446 join two or more observations and 15,365 span more than one day.

Retrieval propagates through identity: a matched observation becomes an entry point that follows same-instance edges to other observations of the same inferred instance and context links to their surrounding episodes and source frames, while the biography excerpt lists not-yet-selected appearances as concrete targets for further search. Retrieval-based methods fetch evidence by semantic similarity, and retrieving relevant events does not necessarily recover the biography of the particular entity a question concerns; GEB establishes identity during memory construction so retrieval can follow an entity across days. On EgoLifeQA, evidence hits rise from 37.6% with MAGIC-Video to 58.9%; on MultiHop-EgoQA, complete evidence coverage rises from 35.4% to 52.1% (a 47% relative gain), with larger relative gains for questions requiring several intervals.

With the same controller, answer model, and retrieval limits, GEB exceeds prior memory frameworks on four benchmarks: EgoLifeQA 72.0% (+4.4 points), Ego-R1-Bench 71.3% (+6.6), MM-Lifelong Test@Week 36.83% (+5.41), and Test@Day 17.58% (+0.83). These comparisons share episodic captions, topic and event summaries, and retrieval limits with MAGIC-Video, so the gains come from memory organization rather than more context; on MultiHop-EgoQA, raising MAGIC-Video's allowance from three to six units per round still leaves its complete coverage below GEB. The 95% confidence intervals exclude zero for EgoLifeQA, Ego-R1-Bench, and Test@Week; on Test@Day the 0.83-point gain over the strongest baseline ReMA includes zero, while the intervals against MAGIC-Video and WorldMM are [+3.34,+11.84] and [+4.83,+13.50].

Ablations attribute the gains to specific parts: removing association costs 3.4 points, keying identity by name costs 2.8, appending descriptions to captions costs 3.8; removing same-instance edges costs 3.0, disconnecting observations from episodes and source clips costs 3.8, and removing both costs 5.0; withholding biography text from both models costs 4.0, from the answer model only 1.2, and removing the unsearched-observation line costs 1.8. These controls show that the benefit of physical-instance organization and biography reading is not explained by additional descriptions or by graph retrieval alone. Ablations run on EgoLifeQA with the controller and answer model fixed, and the same direction reappears on both MM-Lifelong splits (appending descriptions costs 3.25 and 5.08 points on Week and Day).

Perspective

The result targets long-video question answering that requires following the same person or object across events, such as week-long egocentric life recordings and long gameplay streams; the memory is built once offline and serves all later questions, and retrieval is restricted to records preceding the query time, so it suits timestamp-constrained question answering. What a practitioner can reuse directly is the observation-entity-episode-source-clip memory structure, retrieval that propagates along identity, and a controller interface that treats unretrieved appearances as further search targets.

Identity is inferred from visual evidence, and reliable tracking and re-identification across long videos remain open challenges: association errors may assign an observation to the wrong biography, and contextual descriptions may attribute a nearby action to the wrong entity; the larger memory graph also incurs higher latency than caption-only retrieval (the paper reports 5.6 s per search on average and 13.6 s per question). Conservative association can leave one physical instance under multiple identifiers, and the paper uses a note on same-named biographies to distinguish pairs with visual separation evidence from those whose identity remains unresolved. On Test@Day the confidence interval on the gain over the strongest baseline includes zero, so the advantage on that split needs more evidence. In addition, some numbers are missing in the loaded text at the points where they are cited (for example, the raw EgoLifeQA evidence-hit and some accuracy figures are not fully rendered in the abstract and body), so exact reproduction should check the original tables.

Sources