Skip to main content
Back to timeline
arXivSource publication:

QASLR lifts long-video retrieval Hit@1 from 54.2 to 70.7 and holds at both 2B and 8B backbones

Related research and updates

Synopsis

The authors introduce Query-Aware Streaming Latent Reasoning (QASLR), a post-training framework that selects a bounded frame set, restores its temporal order, and processes it clip by clip through a vision-language backbone, where persistent think tokens accumulate cross-clip evidence and an embed token reads out a fixed-dimensional representation after each update; training combines final contrastive learning, step-wise contrastive supervision, and final-embedding self-distillation, raising HourVideo retrieval Hit@1 from 54.2 to 70.7 and from 57.4 to 72.8 on 2B and 8B Qwen3-VL-Embedding, with gains on the evaluated moment-retrieval and video-QA tasks and transfer of the streaming head to the same-family RzenEmbed 7B.

Source-provided article image: Enhancing Long-Video VLM Embeddings with Query-Aware Streaming Latent Reasoning
Figure 1 ·

Figure 1: From one-shot video encoding to streaming evidence accumulation. The upper-left pathway illustrates a one-shot baseline that pools representations from sparsely sampled frames into a single video embedding. The lower-left pathway shows QASLR processing selected frames as ordered clips and producing embeddings throughout the sequence. The right-hand panels summarize example training signals and the intended capabilities of persistent memory, intermediate supervision, optional early exit, and parameter-efficient post-training.

arXiv

Interpretation

QASLR decouples temporal coverage from representation size: a persistent latent memory accumulates evidence clip by clip in chronological order and emits a fixed-dimensional embedding without placing all selected frames in one backbone context. Prior approaches extend context, compress visual tokens, aggregate clip descriptions, or select query-relevant frames; QASLR complements input selection and compression by accumulating selected clip evidence into a fixed-dimensional embedding. The method specifies streaming encoding, memory update, and readout equations; ablations show HourVideo Hit@1 drops by 6.6 points with shuffled clips, 5.7 with reversed clips, 10.4 with memory reset, and 13.8 with only the last clip, and QASLR exceeds the reported Transformer aggregator by 4.5 points.

The training objective supervises intermediate and final representations together: final contrastive learning, step-wise contrastive supervision, and self-distillation with the final embedding as a stop-gradient teacher make partial-video readouts usable for retrieval. Relative to final-only supervision, step-wise supervision plus self-distillation raises the mean Hit@1 over four checkpoints (25%, 50%, 75%, all clips) from 57.3 to 65.1 and the final Hit@1 from 67.9 to 70.7. A four-way loss-combination ablation holds all other settings identical; the 25% checkpoint rises from 43.8 to 55.6 and the final from 67.9 to 70.7, with the authors noting larger gains at the early checkpoint.

The full QASLR recipe improves every reported task at both 2B and 8B scales and transfers to a same-family backbone. Against the original checkpoints, HourVideo retrieval Hit@1 rises by 16.5 and 15.4 points; against single-round post-training, by 15.7 and 15.0 points; on RzenEmbed 7B, QASLR exceeds both baselines on all seven displayed tasks, with gains of 3.4–6.0 points over the original checkpoint. Nine long-video tasks, two model scales, and averages over three random seeds; mean retrieval Hit@1 rises from 45.8 to 53.5 (2B) and 47.2 to 54.8 (8B), and mean QA accuracy from 48.3 to 50.6 and 61.2 to 66.3.

Interleaved sub-batching is a substantial source of the gain and benefits both single-round and streaming post-training. With sub-batching, QASLR mean retrieval Hit@1 rises from 54.5 to 61.1 (2B) and 55.7 to 63.0 (8B); single-round post-training also rises from 52.2 to 56.8 and 53.8 to 59.9. This configuration changes both the effective batch size and the negative pool; under sub-batching QASLR keeps higher Hit@1 on all three retrieval tasks at both scales, while its QA advantage is task-dependent.

Perspective

The work targets retrieval and ranking settings that must compress long videos into fixed-dimensional embeddings: query-independent selection and encoding yield one reusable representation per video for corpus indexing, while the query-aware path yields query-conditioned representations that the authors describe as intended for bounded candidate sets or second-stage reranking. The method is validated on 2B and 8B Qwen3-VL-Embedding and transfers to same-family RzenEmbed 7B; only the streaming head is trained (roughly 200 million and 600 million trainable parameters) with the backbone frozen, and post-training primarily uses LLaVA-Video-178K supplemented with 10,000 ActivityNet videos, with zero source-video overlap between post-training and evaluation sets. Evaluation uses bounded candidate sets (video retrieval generally 4 or 8 candidates, moment retrieval 8 or 16 segments, classification 4–8 categories, QA four options), and results are averaged over three random seeds. The optional early-exit mechanism targets local-clip QA, while fixed-length retrieval uses the complete sequence.

Several open questions remain for a careful reader: the latent memory cannot recover evidence absent from the selected observations, so results depend on the frame selector; against the sub-batched single-round baseline, retrieval advantages are consistent but QA advantages vary by task and scale, with single-round scoring higher on three of four QA tasks at 2B; order and memory perturbations are applied only at inference, which the authors describe as measuring sensitivity to altered inputs and state updates rather than identifying the internal reasoning process; transfer is evaluated only within the Qwen family; warm-cache latency depends on selection, decoding, candidate count, and hardware, and a fixed memory size does not imply a proportional reduction in total computation; seed-averaged scores describe empirical differences without establishing statistical significance; and the early-exit bounds concern agreement with full-step predictions under the stated calibration assumption rather than unconditional ground-truth accuracy. In addition, several equations, loss weights, and some numeric values appear as placeholders in the text, so while the abstract and table values can be checked, equation-level details require the original paper and appendices.

Sources