Skip to main content
Back to timeline
arXivSource publication:

δ-Vision rebuilds layer-wise visual memory with lightweight MLP adapters, cutting FLOPs to 17.90% while keeping every visual token

Synopsis

Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.

AI-generated editorial illustration: Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

Interpretation

The authors find that visual influence on the text stream is concentrated in a low-dimensional subspace: after blocking visual-to-text attention, restoring only a few directions recovers most of the lost accuracy. Prior efficiency work pruned redundant visual tokens; this work reframes the question from which tokens can be discarded to how small the effective dimensionality of visual influence is, and supplies low-rank intervention evidence. Based on low-rank interventions on visual-to-text information flow and spectral analysis of layer-wise visual hidden states: on Qwen3-VL-4B, visual attention outputs need only 38.0–103.5 directions to retain 95% of spectral energy, and 29.6–37.6 directions on LLaVA-1.5-7B, far below the available rank.

Layer-specific visual states are strongly predictable: a separate lightweight MLP trained per layer approximates Transformer-evolved states from the initial visual embeddings with high cosine similarity and low reconstruction error. This suggests preserving useful visual information may not require repeatedly computing full-dimensional visual states at every layer, supporting a shift from which visual evidence to keep to how visual states are constructed. Approximation quality is measured by cosine similarity and reconstruction error, and downstream accuracy remains high when the prediction replaces the true visual evolution path.

The paper proposes δ-Vision: low-rank adapters (an embedding adapter predicting each layer's memory from initial embeddings, and a recurrent adapter updating it layer by layer) construct layer-wise visual memory that is mapped through frozen key/value projections for text queries, removing visual queries, visual attention outputs, and visual feed-forward computation while keeping all visual tokens. Unlike pruning methods that permanently discard visual evidence, δ-Vision keeps every visual token retrievable and compresses only its layer-wise evolution, offering an efficiency axis complementary to sequence-length reduction. On Qwen3-VL-4B the embedding adapter averages 74.4 and the recurrent adapter 75.4, both above the strongest 5%-retention pruning baseline DivPrune at 65.3; a training-objective ablation shows Supervised-KD reaching 74.4 in 34m53s, versus 68.9 for SFT and 74.0 for OPD (4h38m08s).

Across image, multi-image, and video benchmarks, δ-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, and retains roughly 94–96% of vanilla performance across multiple MLLM backbones. This indicates visual-memory prediction is not an artifact of one model or modality but transfers across scales and architectures, including hybrid attention. Single-image average 74.4 over nine benchmarks; multi-image/video average 52.0 after multimodal training, above the strongest 5%-retention pruning baseline at 50.2; cross-backbone averages of 61.9, 78.6, and 71.6 on LLaVA-1.5-7B, Qwen3-VL-30B-A3B, and Qwen3.5-4B; efficiency at 17.90% of uncompressed FLOPs with 1.30x total and 1.50x prefill speedup.

Perspective

The result targets multimodal LLM inference with long visual sequences, especially high-resolution image and video understanding; its audience is researchers and engineering teams who want to cut layer-wise visual computation without discarding visual tokens. The method assumes a frozen vision encoder, multimodal projector, and language-model backbone, training only lightweight visual-memory adapters, with a default configuration of Qwen3-VL-4B plus the embedding adapter at bottleneck rank 128. It can be stacked with token pruning: combined with DART or DivPrune at 50% retention it drops only about 1.2–1.3 points, indicating the two efficiency axes address distinct sources of redundancy. Layer-skipping experiments further show that removing adapters from the first 5 and last 10 layers lowers the average only from 74.4 to 73.3 while FLOPs fall from 17.90% to 15.64%, leaving room to allocate adapter budget per layer.

Sequential low-rank intervention shows visual dependence is highly non-uniform across depth: middle layers are most sensitive on Qwen3-VL-4B, while LLaVA-1.5-7B shows a more distributed, task-dependent pattern; in the hybrid-attention backbone, canceling visual writes to all 24 linear-attention layers has limited effect, whereas blocking visual reads in only FA layers 11 and 15 causes substantial degradation. This model specificity means the optimal allocation of adapter budget may need recalibration per backbone and task. In addition, multi-image and video results depend on an additionally constructed training-data recipe: the single-image-trained adapter scores 40.7 on MuirBench and rises to 45.5 after multimodal training, indicating data composition materially affects long visual-context performance. Larger adapter rank yields only moderate gains, so the trade-off between that benefit and training cost remains to be characterized more finely.

Sources