Public articles linked to the same research event.
arXiv The work introduces Patch-to-Prune (P2P), a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass: validation-guided forward and backward layer sweeps identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance; evaluated on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks with mutually disjoint calibration, validation, and test partitions, P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%, and its layer-wise analysis suggests visual processing is non-uniformly distributed across decoder depth.
The work introduces Patch-to-Prune (P2P), a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass: validation-guided forward and backward layer sweeps identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance; evaluated on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks with mutually disjoint calibration, validation, and test partitions, P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%, and its layer-wise analysis suggests visual processing is non-uniformly distributed across decoder depth.
The work introduces Patch-to-Prune (P2P), a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass: validation-guided forward and backward layer sweeps identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance; evaluated on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks with mutually disjoint calibration, validation, and test partitions, P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%, and its layer-wise analysis suggests visual processing is non-uniformly distributed across decoder depth.
The work introduces Patch-to-Prune (P2P), a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass: validation-guided forward and backward layer sweeps identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance; evaluated on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks with mutually disjoint calibration, validation, and test partitions, P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%, and its layer-wise analysis suggests visual processing is non-uniformly distributed across decoder depth.