FlashGaze prunes video patches before ViT encoding with Quadtree DP, holding 98% of LongVideoBench accuracy at a 20% retained ratio on Qwen3-VL-8B while cutting peak GPU memory by 1.8x
Synopsis
FlashGaze is a training-free method that reduces spatiotemporal redundancy before ViT encoding by using pixel-space differences as a proxy for information loss and Quadtree Dynamic Programming to jointly decide patch dropping, merging, and keeping under a fixed budget; on Qwen3-VL-8B and NVILA-8B-Video it achieves higher accuracy and lower TTFT across multiple video understanding benchmarks, and at a 20% retained ratio preserves 98% of the full-input LongVideoBench accuracy.
Interpretation
Redundancy reduction is moved to the video input stage, so patch selection happens before ViT encoding and dropped redundant content never enters the encoder. Prior methods prune inside or after the ViT (ToMe, VisionZip, FastV), leaving encoding cost unaddressed; AutoGaze selects patches before encoding but relies on an autoregressive network and specially constructed training data. FlashGaze introduces no auxiliary network and requires no training. Evaluated on two backbones, Qwen3-VL-8B and NVILA-8B-Video, across six benchmarks; the stage-wise latency breakdown in Appendix D shows FlashGaze pre-ViT overhead around 0.5 s versus 7.3 to 14.6 s for AutoGaze.
Pixel-space differences measure both temporal redundancy (current patch versus its historical patch in visual memory) and spatial redundancy (reconstruction difference before and after merging), and Quadtree DP jointly optimizes dropping, merging, and keeping under a budget. Multi-scale patch selection is cast as a budget-constrained objective with a patch-count penalty, solved over a three-level fine-mid-coarse quadtree with Subtree Cost Merging, which avoids the conflict of selecting overlapping candidates at different scales independently. Ablation at a 30% retained ratio with 768 frames shows pixel differences reach 71.7 on HLVid and 66.3 on LongVideoBench, better than patch embedding, SigLIP, and LPIPS; Quadtree DP attains the lowest cumulative action cost of 1.96 versus random, single-scale greedy, and multi-scale greedy, with the highest accuracy of 71.7.
At matched retained ratios it delivers both higher accuracy and lower latency and memory, and it scales the backbone to longer, higher-resolution inputs. On LongVideoBench and HLVid, FlashGaze records the lowest TTFT and lowest peak GPU memory at 50%, 40%, 30%, and 20% retained ratios; at 20% it preserves 98% of full-input LongVideoBench accuracy, reduces peak GPU memory by 1.8x, and achieves a 13x TTFT speedup over the vanilla model, including 5.4x in ViT encoding and 17x in MLLM prefill. Tables 1 and 3 report accuracy, peak GPU memory, and TTFT across two backbones and several retained ratios; Table 2 shows that scaling Qwen3-VL-8B to 1536 frames and 3840 resolution yields higher accuracy than the vanilla model on all six benchmarks.
Analysis indicates pruned visual representations remain highly similar to the full input, the model allocates more attention to dynamic visual tokens, and fine-scale patches land in information-rich regions. Representation similarity is measured on 20 HLVid videos by using the last prefill token as the query and aggregating first-layer visual value vectors by attention weights, then computing cosine similarity; dynamic-token attention is measured at a 50% retained ratio. Attention analysis shows FlashGaze consistently assigns more attention to dynamic tokens than the vanilla model at each layer while HLVid accuracy rises from 71.6% to 73.2%; visualization shows retained ratios of only 10% and 13% on mostly static clips, with coarse-scale patches for low-information areas such as sky and fine-scale patches for detail-rich regions such as traffic signs.
Perspective
The method targets spatiotemporal redundancy at the video input stage and applies to MLLM backbones that support multi-scale patch inputs: the Qwen3-VL series natively supports multi-scale image inputs and integrates seamlessly without downstream fine-tuning, while NVILA-8B-Video enables multi-scale patch inputs by interpolating each frame and its positional embeddings to different scales. Evaluation covers six benchmarks spanning general, long-video, and high-resolution settings, with main comparisons on LongVideoBench, VideoMMMU, and HLVid, and comparisons of accuracy, peak GPU memory, and TTFT under different retained ratios on Qwen3-VL-8B. Latency is measured on a single NVIDIA H200 GPU with batch size 1 and BF16 precision, consistently using PyTorch native SDPA rather than FlashAttention, with one warm-up run followed by three consecutive runs averaged. For settings that need low-latency processing of long, high-resolution video, this means trading the same hardware for longer context and higher resolution.
How the per-frame Quadtree DP solve cost scales with frame count and resolution is not expanded; the text reports only a combined pre-ViT latency of about 0.5 s. The behavior of pixel differences as an information-loss proxy under strong camera motion or global illumination change is not reported, since the text mentions only block-wise local search to mitigate slight camera motion. Visual memory is updated with the most recently retained patch, and how content reappearing after long occlusion is handled is not discussed. The attention analysis rests on 20 HLVid videos at a 50% retained ratio, and the link between higher dynamic-token attention and accuracy gains is framed as a possible contributor rather than an established causal mechanism. In addition, several speedup factors appear blank in the abstract and conclusion, so concrete values should be read from Table 3 and the stage-wise latency tables in Appendix D.
