Skip to main content
Back to timeline
arXivSource publication:

APM-Bench tests cross-session persistent memory for streaming video assistants across 549 sessions and 104 life trajectories, finding existing methods struggle to combine recall, latency, and storage

Synopsis

The work introduces APM-Bench, which reformulates egocentric streaming interaction as multi-session life trajectories (549 sessions, 104 trajectories, 2,719 candidates, averaging 69 minutes of video per trajectory) and evaluates general video models and eight specialized memory systems on cross-session understanding, real-time perception, and adaptive response, plus an evidence-availability-aware test; results reveal a clear utility-latency-storage trade-off: raw video memory gives the strongest cross-session performance but needs GiB-scale storage and high latency, text summaries cut storage to the KiB scale at the cost of cross-session performance, event-structured memory shows the strongest utility among specialized systems, and adaptive response remains difficult even with rich history

AI-generated editorial illustration: APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

Interpretation

APM-Bench reformulates real-world egocentric streaming interaction as multi-session life trajectories that preserve realistic time gaps between sessions, turning "can memory persist across interruptions" into a measurable evaluation target. Prior streaming video benchmarks mostly focus on a single continuous video or short clips; APM-Bench uses 549 sessions, 104 trajectories, and 2,719 candidates across 12 tasks in three capability families, preserving continuous temporal order within sessions and realistic gaps between them. The benchmark builds on two egocentric datasets, EgoLife and HD-EPIC, through a two-stage pipeline with automated and human review; two annotators reach Cohen's kappa agreement on 300 sampled questions, and each trajectory averages 69 minutes of video with individual sessions averaging 13 minutes.

Evaluation shows a clear utility-latency-storage trade-off across memory representations: raw video memory is strongest on cross-session understanding, text summaries cut storage to the KiB scale but lose cross-session performance, and event-structured memory shows the strongest utility among specialized systems. Prior evaluations rarely assess memory utility together with deployment costs; by organizing intermittent interactions into sessions, APM-Bench defines memory-storage boundaries and enables comparison of utility, storage, and first-token latency across representations on the same trajectories. General video models and eight specialized systems are evaluated under uniform 1 FPS input and the same host configuration; raw video memory reaches GiB-scale storage with high query-time overhead, text summaries are KiB-scale, event-structured methods lead among specialized systems, and parametric memory keeps state from growing with video length but shows limited utility and nontrivial latency.

Excessive history can interfere with current-scene understanding, so persistent memory should be injected selectively rather than retained in full. This separates "what to store" from "when to use it": for most general video models, replacing raw-video history with text summaries improves real-time perception, indicating that information useful for later recall may be unnecessary or distracting in the current interaction. Real-time perception tasks are evaluated on the current session's causal video prefix, comparing the same models under raw-video memory and text-summary memory; in the appendix, FluxMem's real-time perception score falls from 33.95 under Oracle to 25.14 with APM-Bench and 15.28 with Raw Lifelong as supplied video grows.

Strong recall does not imply awareness of missing evidence, and adaptive response remains difficult even with rich history. APM-Bench builds a dedicated evidence-availability-aware set and distinguishes assistance driven by explicit registrations (RCR, PRM) from fully autonomous proactive assistance (MPA, TPG), separating "when to speak" from "whether the response is right." The evidence-availability-aware set contains 260 cross-session understanding questions where models access only the two most recent completed sessions, with 130 questions answerable from that history and 130 requiring evidence outside it; Gemini 3.6 Flash is strong when evidence is accessible (85.38%) but weaker at recognizing unavailability (59.23%), while SimpleStream shows the opposite pattern; on adaptive response, even with raw video memory the strongest general models remain weak on MPA and TPG relative to RCR and PRM.

Perspective

The benchmark targets egocentric streaming video assistant settings where past visual experience must be reused after interruptions while storage and response latency stay manageable; its trajectories come from EgoLife and HD-EPIC, covering everyday activities and structured kitchen procedures, so conclusions apply most directly to wearable or mobile personal assistants. For researchers comparing memory representations (raw video, text summaries, KV cache, visual tokens, event trees, parametric memory, reasoning thoughts), APM-Bench offers a testbed that observes utility, storage, and first-token latency on the same trajectories, and its evidence-availability-aware set provides a measurable target for when an assistant should acknowledge insufficient evidence.

Evaluation runs under uniform 1 FPS input and the same host configuration, and storage for specialized systems is estimated from the memory state maintained during inference because most methods target single continuous videos and do not export persistent state; how this estimation affects conclusions deserves attention. The evidence-availability-aware set grants access only to the two most recent completed sessions, so whether its conclusions hold under longer accessible history remains open. Adaptive response is scored by an LLM judge on open-ended answers, and agreement between the judging rubric and human judgment warrants further observation. In addition, the comparison between trajectory organization and longer raw history covers only 12 EgoLife trajectories and 366 candidates, a limited sample.

Sources