Skip to main content
Back to timeline
arXivSource publication:

V-CoLA compresses vision tokens using linear-attention intrinsic signals, retaining 88% performance at 12.5% tokens with up to 6.15x prefill speedup

Synopsis

The authors propose V-CoLA, a training-free vision token compression framework for hybrid vision-language models with linear attention, deriving a uniqueness-aware importance criterion from the retention preferences of the bounded recurrent state, combined with adaptive chunk-wise merging and early exiting of deep-layer vision tokens; on Qwen3.5-9B and related models it reaches 99.5% of original performance at 50% vision tokens and over 88.0% at 12.5%, with 1.86 to 6.15 prefill speedup.

Source-provided article image: V-CoLA: Vision Token Compression with Linear Attention
Figure 1 ·

Figure 1: Comparison of token selection. FastV suffers from attention bias, missing the digits at the top of the image. DART selects pivot tokens from only sparse semantic regions, leading to insufficient awareness of some primary objects. Our method successfully identifies all main regions and provides an accurate response.

arXiv

Interpretation

The paper systematically analyzes the structural failure of existing vision token compression methods on hybrid architectures with linear attention: both attention-based and similarity-based methods show virtually no advantage over a random pruning baseline after transfer. Prior methods rely on softmax attention scores or latent feature similarity and are mostly evaluated on softmax-attention models; this work shows that shallow-layer attention concentration and latent feature discriminability are both altered in hybrid architectures, requiring new importance signals. Qualitative evidence from multi-benchmark comparisons (Figure 2), attention distribution comparisons across depths in Qwen3-VL and Qwen3.5 (Figure 3), and analysis of DART redundant token selection (Figure 4).

The paper proposes a uniqueness-aware importance criterion that jointly weights each token's long-term contribution to the final state and its short-term uniqueness relative to historical context, identifying critical vision tokens while avoiding redundancy. Compared with using only the magnitude of state update, this criterion accounts for both long-term information retention and short-term redundancy suppression, defining importance from the reconstruction objective of linear attention. Constructed from reconstruction error and a pseudo-query probe of state change; the weight parameter is compared in ablations, and the default setting yields the best results across model scales and architecture families.

The paper proposes an adaptive chunk-wise merging strategy that partitions tokens into importance-balanced chunks according to importance distribution density and merges within each chunk, giving finer granularity to important regions. Unlike clustering-based or bipartite-matching aggregation, this strategy explicitly considers the importance distribution, adaptively allocating compression rates across semantic regions. Ablations show that removing adaptive merging causes performance drops, and tuning the temperature coefficient can emulate hard pruning, offering flexibility.

The paper implements optimizations compatible with the chunk-wise parallelism of linear attention and adds an early-exit mechanism for deep-layer vision tokens, substantially reducing prefill overhead while preserving multimodal reasoning performance. It extends Gated DeltaRule to support multiple query injection so pseudo-query probing works under chunk-wise parallelism; early exit reallocates the deep-layer redundant token budget toward shallow-layer visual understanding. Adaptive token merging adds only 0.86ms, about 0.1% of total prefill cost; the extended Gated DeltaRule achieves a 3.85x speedup under 4-query inference; early exit is validated on V*-Bench for the value of intermediate-layer visual access.

Perspective

The result targets hybrid vision-language models that interleave linear-attention and softmax-attention layers, such as Qwen3.5 and InfiniteVL, and suits multimodal inference settings where reducing vision token overhead in the prefill stage matters. The method is training-free and can be layered onto existing models, with a default configuration achieving the best results on Qwen3.5-9B, Qwen3.5-27B, and InfiniteVL without per-model tuning. For recurrent backbones with compatible matrix-state readouts, both criteria remain well defined, and extending evaluation to such backbones is a natural next direction.

Evaluation mainly covers hybrid vision-language models interleaving linear and softmax attention; purely Mamba-based or fully linear-attention vision-language models are not yet included, and extending the method to such architectures may require revisiting how the uniqueness-aware criterion interacts with the absence of periodic softmax-attention layers. The analysis and method design are mainly grounded in the Gated DeltaNet formulation, and a more systematic study across a broader range of linear-attention formulations would help assess the generality of the criterion. In addition, some equations and figure values appear as placeholders in the text, so reproducing specific coefficients and curve details would require consulting the original figures and tables.

Sources