P2P turns activation patching into an inference-time bypass, keeping about 94% dense accuracy at 3% tolerance while cutting FLOPs by 55%
Related research and updatesSynopsis
The work introduces Patch-to-Prune (P2P), a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass: validation-guided forward and backward layer sweeps identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance; evaluated on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks with mutually disjoint calibration, validation, and test partitions, P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%, and its layer-wise analysis suggests visual processing is non-uniformly distributed across decoder depth.
Figure 1: Framework of Patch-to-Prune
arXivInterpretation
Introduces P2P, a training-free framework that turns activation patching from a diagnostic into an inference-time computation bypass, replacing visual-token projection outputs in selected decoder regions with fixed neutral proxy activation vectors. Unlike conventional token pruning, P2P preserves sequence length, token order, positional information, attention mask, and residual pathways, so it prunes computation without removing tokens or modifying pretrained model weights. The method description specifies forward and backward layer sweeps, validation guidance, and a user-specified accuracy tolerance; evaluation covers four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks using mutually disjoint calibration, validation, and test partitions.
At a 3% tolerance, P2P retains around 94% of dense accuracy while reducing FLOPs by 55%. The result reports accuracy loss and compute savings under the same tolerance parameter, making the fidelity-efficiency trade-off explicitly tunable rather than fixed by a pruning ratio. The numbers come from the multi-model, multi-benchmark evaluation described in the abstract; per-model and per-benchmark breakdowns are not given in the loaded text.
Layer-wise analysis suggests visual processing in VLMs is non-uniformly distributed across decoder depth: early and late layers often require little token-specific visual computation, whereas intermediate layers appear to perform most task-relevant visual integration. This observation uses the efficiency framework as a causal lens, indicating that later reasoning can rely largely on visual information already embedded in shared residual and textual representations. The conclusion is phrased as "suggests" and "appear to," making it a causal hint from layer sweeps rather than an established mechanistic finding.
Perspective
The framework targets engineering and research settings that aim to reduce VLM inference compute without retraining or altering pretrained weights, tuning the extent of the computation bypass through a user-specified accuracy tolerance. Its evaluated scope is four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks as described in the loaded text; the non-uniform depth observation offers a testable hypothesis that intermediate layers carry most visual integration while early and late layers need little token-specific visual computation.
The loaded text is abstract-level and contains no figures or per-benchmark results, so the variability of the 94% accuracy retention across models and task types cannot be judged, nor can 3% be confirmed as the optimal operating point. The non-uniform depth observation is phrased as "suggests" and "appear to," a causal hint rather than a settled finding, and its stability across VLM architectures and visual tasks remains an open question. In addition, how the neutral proxy activation vectors are constructed and the search cost of the validation-guided sweeps are not described in the loaded text.
