LEAP keeps hour-scale audio-video QA out of one giant context: 4.5–16.8% over the Qwen3-Omni-30B baseline and transfer to MiniCPM-o 4.5
Synopsis
LEAP splits an hour-scale recording into fixed-duration blocks, runs a lightweight localization pass per block to score short candidate windows, then re-encodes only the top-ranked windows in a single bounded answer pass, so the answer input and peak context stay independent of recording duration; across several AVQA benchmarks it improves over the Qwen3-Omni-30B-A3B baseline by 4.5–16.8% and transfers to MiniCPM-o 4.5, surpassing its published results by 3.1–13.0%.
Interpretation
LEAP introduces a two-stage framework in which the model retrieves its own evidence: a recording is partitioned into non-overlapping fixed-duration blocks, each block into short candidate windows, a lightweight question-conditioned localization pass scores the windows, and the highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Prior selection methods pick regions from a coarse view, compression methods keep the full recording but dilute detail, and agentic methods leave where to look to general-purpose models prompted at inference; LEAP decouples evidence localization from reasoning and trains both stages, with a localization LoRA improving the selected windows and an answer LoRA improving the answers read from those same windows. The paper gives a formal two-stage formulation, the option-letter logit scoring rule and training objective for the localization pass, and deployed configuration values for block length, candidate windows, retained blocks and windows; ablations show that removing retrieval, the localization LoRA, or answer training lowers accuracy.
Because each pass reads one block or a bounded set of retained windows, the answer input, peak context and peak memory are independent of recording duration; the paper provides token accounting showing each pass has a bounded context. A whole-recording read costs a large number of tokens per hour and exhausts the backbone's position limit at roughly tens of minutes; LEAP reads one block or at most a small set of retained windows per pass, so cost does not grow with duration. The paper gives a cost equation and deployed figures for localization-pass and answer-pass token caps and the answer-pass position cap, and states that in deployed runs every block pass and answer pass stays under those caps.
The same block grid can localize windows over pre-computed transcripts without decoding media frames, while the final answer pass still reads the raw audio-visual stream, preserving visual and non-speech evidence a transcript misses. Existing caption-based or document-retrieval pipelines often hand their answering model text alone; LEAP lets the transcript channel and the media channel share the block grid and the answer pass, differing only in what is scanned. The paper reports that the transcript channel is close to the media channel in accuracy on several benchmarks while spending fewer localization tokens and answering faster per question, and that the more of the evidence is spoken, the more of it the transcript channel retains relative to the media channel.
The block grid natively supports causal queries, letting LEAP support streaming inference without streaming-specific training; on StreamArena's historical-retrospection task LEAP significantly leads both a whole-prefix read and uniformly spread windows on the same backbone. Streaming systems are usually purpose-built around a KV cache, a fixed-size memory, or a textual memory of distant history; LEAP reuses the clock-fixed block grid, with the localization pass selecting windows only over blocks up to query time and the answer pass re-reading them from raw media. The paper reports that on StreamArena both LEAP's media and transcript channels significantly lead the whole-prefix and uniform-window baselines, and that on the 9B MiniCPM-o 4.5 LEAP significantly beats that model's native streaming mode.
Perspective
This work targets hour-scale audio-visual question answering, especially settings where evidence is scattered across a few brief windows and where speech and non-speech or visual cues both matter; it applies to a deployment where backbone weights stay frozen and only two LoRA adapters are trained, and to a causal streaming setting where only the recording prefix up to query time is readable. For a reader, it means a long recording need not enter the context whole: the localization pass covers every block, the answer pass spends its context on a few retained windows, and each window is read at a per-window density that does not change with duration; the transcript channel is a cheaper substitute where the evidence is spoken. The paper also names open directions, including stronger retrieval scores within the bounded-cost structure and an answer pass that can tell when its retained windows miss the evidence and return to the scan for more.
The localization adapter is trained on LongVALE-derived questions whose clips all fit inside one block, so block length is set by the longest training clip; how this behaves on longer recordings or on evidence that spans blocks more intricately is something a reader can keep watching. The transcript channel retains less evidence than the media channel when the evidence is primarily visual, as the paper reports on CG-Bench mini; so the boundary for substituting a transcript scan for a media scan depends on how much of a recording's evidence is spoken. The streaming results come from StreamArena's historical-retrospection task, and under the causal-access protocol the answer pass uses base weights with no adapter mounted; more general streaming interaction remains to be observed. The paper also states that code will be released upon acceptance, so present reproducibility rests on the configuration details in the main text and appendices.
