LT-OPD uses on-policy self-distillation on its own generated trajectories to lift average retained performance at 5% visual tokens from 68.6% to 82.3%
Synopsis
The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
Interpretation
The authors reframe extreme visual-token reduction as a training problem: a low budget not only removes visual evidence but also changes the generation states the compressed model encounters at inference, so training on fixed prefixes does not cover the states the student itself predicts. Prior work splits into training-free selection methods (FastV, DivPrune, DART, CDPruner, HiPrune, ZOO-Prune) and fixed-prefix training-based adaptation (MQT-LLaVA, LLaVA-Mini, EPIC, LearnPruner); the paper uses a comparison table that separates methods by prefix/state source and on-policy status, noting that none of them learn on their own generated states. The paper motivates the method with a figure showing that under a low budget the compressed model's predictive distribution deviates from its full-token counterpart; this is the authors' own empirical motivation and no independent quantitative deviation measure is reported.
LT-OPD lets a low-token student roll out its own responses while a frozen full-token copy of the same model provides next-token distributional supervision on those student-generated prefixes, using Jensen-Shannon divergence with the sampled trajectory treated as stop-gradient. Unlike progressive consistency distillation such as EPIC, or on-policy distillation aimed at stronger or more informative teachers, the teacher here is a frozen full-token copy of the same model and the supervision positions are chosen by the compressed student itself. The method section gives the problem formulation, rollout sampling, teacher and student distribution definitions, and the JSD objective; implementation details (top-100 vocabulary support plus one aggregated tail category, mixture coefficient 0.5, loss averaged over valid response tokens within a sequence then across sequences) appear in Appendix B.
A budget-level curriculum decays the retention ratio along a cosine curve down to 5%; ablating the final hold at 5% drops average retained performance from 82.3% to 79.5%, while fixed-prefix JSD knowledge distillation reaches 77.6%. The paper separates the curriculum from on-policy learning in ablation: the curriculum adds 2.7 percentage points over plain SFT, and on-policy distillation alone without curriculum reaches 81.5%, indicating both components contribute. Ablations run under the same training configuration, initialization seed, and 175-update budget, and are evaluated at 5% retention under a unified evaluation protocol.
On nine benchmarks with Qwen3.5-4B at 5% visual tokens, average retained performance rises from 68.6% to 82.3%, above the strongest training-free baseline DivPrune at 77.6% and the training-based baseline EPIC at 73.7%, and is best among prior methods at the same budget on 8 of 9 benchmarks. The paper highlights that at a 5% budget the result exceeds all training-free methods retaining 10% of tokens (at most 80.2%) and nearly matches the best results at 15% or 20% budgets, suggesting adaptation training can partly substitute for keeping more tokens. Results come from full tables across nine benchmarks spanning perception-heavy (V*Bench, HR-Bench 4K), general VQA (GQA, MMMU, MMB, MME), hallucination (POPE), and OCR (TextVQA, OCRBench) categories; a comparison with GRPO, GSPO, and DAPO shows LT-OPD's 82.3% above their 75.8%-78.2%.
Perspective
The result targets deployment settings that must run multimodal large language models under extreme visual-token budgets, such as edge devices or resource-constrained inference; the stated conditions are 5% visual-token retention, CDPruner as the selector, and training on LT-14K (14,000 single-image questions). Methodologically it applies to any MLLM that can supply a frozen full-token copy of itself as teacher, which the paper verifies on Qwen3.5-4B/9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. Efficiency gains concentrate on the prefill side: KV cache and prefill FLOPs each drop by about 85%, while end-to-end latency falls only modestly from 1331.5 ms to 1293.7 ms because autoregressive decoding time is dominated by per-token computation and memory access. Training cost is 15.1 hours on eight A100 80GB GPUs, 175 updates, and 112,000 generated trajectories.
The paper and appendices contain many tables, but some content (full explanations of different training objectives, multi-image benchmark details, qualitative cases) sits in appendices, so reading only the abstract or main text makes it hard to fully judge the relative contribution of each component. The teacher is always a frozen initial full-token policy; the paper also tests trust-region regularization and an EMA teacher, with average retained performance between 81.3% and 82.3%, indicating insensitivity to teacher construction but leaving the potential benefit of a teacher that evolves during training largely unexplored. Objective ablations show JSD is better on perception-heavy benchmarks while forward and reverse KL score higher on OCR benchmarks, suggesting the best objective may depend on task type. Multi-image and video results are listed by the authors as a future direction, and behavior in video understanding still needs more validation. Efficiency measurements are based on a fixed reference replay over 305 TextVQA questions, so real deployment latency may vary with load and hardware.
