SCOPD lifts VLM performance at 10% visual-token retention from 86.37% to 92.43% via sparse-context on-policy self-distillation
Synopsis
Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
Interpretation
The paper identifies and tests a representation-utilization gap: on 500 image-question pairs, the visual representation is pruned once and held fixed, greedy decoding drops sharply, yet repeated sampling from the same sparse representation recovers many otherwise failed examples, while a language-only control stays near zero. Prior work commonly attributed sharp degradation under aggressive pruning to irreversible loss of task-relevant visual information; this analysis indicates the evidence can remain accessible but is not reliably used during reasoning. Fixed-context Pass@64 diagnostic on 500 free-form numerical image-question pairs (from MMStar, CVBench, MMMU-Pro, BLINK, LogicVista, LLaVA-CoT) with VisionZip at 10% retention, using GPT-6-Astra-Max as a reasoning-validity judge and a language-only control near zero.
SCOPD has a sparse-context student generate reasoning trajectories while a privileged full-context teacher supervises the same on-policy prefixes at token level, requiring no ground-truth answers, reasoning traces, or architectural changes. It moves on-policy self-distillation from language-model or model-structure pruning settings to the case where the conditioning representation itself is deliberately sparsified, using the corresponding unpruned representation as privileged supervision. Across 13 image benchmarks, normalized average rises from 86.37% (Vanilla) to 90.49% at 10% retention and from 93.42% to 96.80% at 20% retention, outperforming SFT, GRPO, and EPIC baselines.
SCOPD+ uses a small visual-budget intervention to identify response positions sensitive to visual evidence, distilling only those positions and backpropagating through only a fraction of response tokens. Rather than relying on teacher-student KL divergence alone, it perturbs the student's available visual evidence while holding the reasoning prefix fixed, separating visually relevant disagreement from ordinary language-level variation. At 10% retention Avg13 reaches 92.43%, improving over SCOPD on 8 of 13 benchmarks and matching it on two; in ablations SCOPD+ reaches Avg6 95.25 versus dense SCOPD 94.44, Top-KL 94.51, TIP 94.21, random 93.56, while selecting the least sensitive positions drops to 89.13.
Sparse-context adaptation transfers across pruning operators, model families, and image-to-video evaluation without adding inference cost. Training uses only VisionZip, yet evaluation without further adaptation improves under DivPrune, random pruning, and FastV; Qwen3-VL-4B and five video benchmarks also benefit. Avg6 improves under every tested pruner (VisionZip 88.74 to 95.25, DivPrune 83.47 to 91.23, random 82.06 to 88.85, FastV 81.47 to 86.03); on Qwen3-VL-4B 75.29 to 81.60/82.92; video Avg5 from 90.88 to 95.70/96.60; SCOPD+ adds 21.4% theoretical training compute but only 1.9% measured time, with no extra inference-time computation.
Perspective
The result targets reasoning VLMs that use training-free visual-token pruning at deployment: when task-relevant evidence remains in the sparse representation, SCOPD and SCOPD+ use on-policy self-distillation to help the language model use that evidence more reliably, under fixed retention budgets for image and short-video reasoning. Training needs only image-question pairs, no ground-truth answers or reasoning traces, with the vision encoder and multimodal projector frozen and only LoRA adapters on the language model; it is therefore a post-training step on top of existing pruning pipelines rather than a replacement for the pruning operator itself.
The fixed-context Pass@ diagnostic is a controlled diagnostic set (500 free-form numerical image-question pairs) rather than official benchmark evaluation, so how broadly its conclusion holds awaits wider task coverage; training focuses on VisionZip with fixed retention budgets, leaving broader in-LLM pruning, dynamic pruning, and jointly learned compression policies open; evaluation centers on short-form image and video reasoning, so whether long-horizon video, embodied, and agentic tasks benefit similarly in retaining and using visual evidence remains an open question; and SCOPD+'s gains depend on the selection fraction and intervention budget hyperparameters, whose optima across models and tasks need further observation.
