Skip to main content
Back to timeline
arXivSource publication:

TRACE audits streaming video understanding across 1,240 records from 517 videos, finding that near-identical accuracy hides large gaps in completion, invalid output, and response timing

Synopsis

TRACE introduces a condition-aware benchmark and evaluation framework for streaming video understanding that makes temporal validity, execution conditions, and operational outcomes explicit, evaluating eight publicly available models or systems in eight configurations on 1,240 records from 517 videos, and finds that nearly identical QA accuracy can mask substantial differences in completion, answer validity, and generation workload, while proactive performance separates into response quality, response delay, false alarms, and missed target windows.

AI-generated editorial illustration: TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

Interpretation

TRACE reframes streaming-video evaluation as a system-level measurement problem in which task scores must be interpreted under explicit execution conditions, recording how visual state is maintained, how responses are initiated, and what the evaluation boundary includes rather than treating these choices as implicit model properties. Earlier streaming and online video benchmarks such as StreamingBench, OVO-Bench, OVBench, and RIVER mainly establish causal visibility, meaning future information must be hidden; TRACE additionally requires declaring when evidence becomes valid, how history is maintained, and who triggers a response. The paper implements this through a unified Core-Adapter protocol: the Core delivers timestamped RGB frames at 1 FPS with real-time pacing and controls question or instruction arrival, while Adapters record actual image submissions, history replay, frame drops, state maintenance, outputs, and failures; each configuration's state organization, proactive trigger mode, and deployment/evaluation boundary is listed in Table 3.

Temporal audit is not a data-cleaning step but the core of the evaluation boundary: in the release set, 328 records were edited and 52 excluded, with 325 of the 328 edited records involving timing or trigger fields. Prior benchmarks typically treat timestamps as fixed metadata, whereas TRACE makes temporal validity an explicit, reviewed part of the evaluation protocol. For directly comparable numeric timestamp fields, the median absolute revision is 3.0 s for QA (127 field changes, 90th percentile 58.2 s) and 2.24 s for Proactive Response (150 field changes, 90th percentile 20.76 s); 100 of 103 edited QA records contained timing revisions, and all 225 edited Proactive records revised or added trigger annotations.

Similar task scores can hide different system behaviors, so TRACE reports quality, timing, response selection, workload, completion, and reliability separately instead of collapsing them into a single aggregate score. The paper uses paired cases to show this separation: LiveCC and MOSS-Preview achieve nearly identical QA accuracy (65.19% and 65.07%) yet differ substantially in completion (93.88% and 100.00%), recorded output tokens (146,064 and 42,206), and invalid-output rate (6.12% and 1.92%); MOSS-VL and AURA have nearly the same In-window Accuracy (8.05% and 7.92%) yet differ in False-alarm Rate (41.1% and 59.1%) and median Response Delay (0.33 s and 1.15 s). Results come from standard-set evaluation of eight publicly available models or systems in eight configurations, with failed runs retained in the denominators; on the shared-video population LiveCC exceeds AURA by 4.92 points (95% CI 2.28 to 7.97) while missing only 0.31% of target windows yet producing false alarms in 65.0% of its assembled response episodes.

Execution conditions define what a score measures: different history organization and system boundaries mean the same numeric metric does not imply the same components, workload, or failure sources. The paper notes that AURA reconstructs the legal visual prefix at QA query time but maintains persistent incremental state for Proactive Response, so its QA latency and submitted-image count describe a different history-processing path; JoyAI's 17.08% In-window Accuracy is measured at a complete-system boundary that includes tested memory and scheduling components. MiniCPM-O native duplex provides an interface-level diagnostic: 93.52% of its spoken-style QA outputs cannot be recovered as a unique option, yielding 2.40% QA accuracy, and on Proactive Response it reaches 0.51% In-window Accuracy with a 65.43% Miss Rate, reported as measurements of the tested interface rather than interface-independent capability estimates.

Perspective

The work targets researchers and system developers studying streaming video understanding evaluation, and applies to visual-only, single-instruction, 1 FPS causal-access settings: QA tasks introduce a question at a specified video time, Proactive Response tasks provide a monitoring instruction before target events or states, and the standard proactive configurations self-initiate responses after one monitoring instruction. It lets a score be read together with temporal validity, execution conditions, and operational outcomes, and provides a reproducible evaluation population and revision provenance; the paper also names natural extensions including no-trigger and negative-event coverage, higher visual sampling rates, audio and multi-turn interaction, and sustained-operation tests.

Several open questions remain for a careful reader: current findings are limited to the tested models or systems on visual-only, single-instruction tasks at 1 FPS; proactive tasks emphasize positive triggers, so the False-alarm Rate measures responses emitted when no response window is currently valid and a later opportunity remains, rather than a general false-positive rate on no-trigger videos; Median Response Delay is computed only over answered windows with an observed onset, and observed-onset coverage ranges from 6.8% to 70.5% across configurations, which limits how broadly that median characterizes all answered windows; and the semantic Judge calibration rests on a stratified sample of 36 responses covering 31 distinct semantic items, with 83.3% human-Judge binary agreement (Cohen's kappa 0.675), which the paper describes as a preliminary human comparison rather than complete ground truth, and the Judge sees no video frames, so it cannot detect visually contradicted answers that appear textually correct.

Sources