Skip to main content
Back to timeline
Hugging FaceSource publication:

Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

Synopsis

This work reformulates block removal in large language models as a constrained binary optimization equivalent to finding low-energy states of an Ising glass, using a once-computed approximate-Hessian energy as a cheap proxy for downstream quality so that vast numbers of candidate configurations can be ranked without benchmarking any of them; on Llama-3.3-70B-Instruct at 50% compression (40 of 80 blocks removed) without retraining, MMLU stays near 77 while the strongest baseline, block influence, falls to about 54, with transfer also shown on Llama-3.1-8B-Instruct, Qwen3-14B, and the hybrid NVIDIA-Nemotron-3-Nano-30B-A3B-FP8.

AI-generated editorial illustration: Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

Interpretation

Block removal selection is formalized as constrained binary optimization, physically equivalent to an all-to-all coupled Ising glass with conserved magnetization, where each binary variable marks a transformer block as kept or removed. Existing block-removal methods mostly score each block independently (magnitude, sensitivity, block influence heuristics), which in physics terms is a mean-field approximation, and often remove only a single consecutive run of blocks; this work uses a second-order Taylor expansion to obtain an approximate Hessian whose off-diagonal entries explicitly capture pairwise couplings between blocks. The method section lays out the chain from second-order Taylor expansion of the loss to the Hessian, then to the energy xᵀH⁰x and the QUBO mapping, with results across several models; the source is an article-style overview, with full derivations and ablations in the paper.

The Ising energy is a strong yet cheap proxy for downstream benchmark quality: the Hessian is computed once from forward and backward passes on a small calibration dataset, after which evaluating any candidate configuration is a single cheap energy calculation, and the same Hessian can be reused for different compression targets M. Compared with approaches that must run or benchmark each candidate configuration, this proxy reduces candidate ranking to a single energy evaluation, making brute-force enumeration of tens of billions of spin configurations feasible on a single GPU. Reports brute-force enumeration on a single GPU up to tens of billions of configurations, with the hardest tractable case, removing 8 of Llama-3.3-70B's 80 blocks (about 29 billion configurations), taking roughly two days; beyond that, QUBO solvers are used.

In the deep-compression regime, coupling-aware CBO substantially outperforms the block influence baseline: for Llama-3.3-70B-Instruct without retraining, removing 32/80 and 40/80 blocks gives CBO MMLU of 76.6 and 76.9 versus 59.3 and 54.0 for the baseline, an advantage of almost 23 points at the deepest setting; for Qwen3-14B at 12/40 removed, CBO leads MMLU by about 10 points. At lighter compression the two are comparable, and the gap widens as compression deepens, indicating that inter-block couplings matter most when many blocks are removed, precisely where mean-field-style independent scoring leaves quality on the table. Provides a concrete MMLU table for Llama-3.3-70B-Instruct without retraining (original 82.2) and states that at the deepest setting CBO beats the baseline on every benchmark tested; the source is an article-style overview, with full result tables in the paper.

The best pruning is often an excited state rather than the ground state: for Llama-3.1-8B-Instruct at 16/32 blocks removed, the 17th excited state is the first to propose removing a block near the beginning of the model, and after light retraining that configuration outperforms the ground state across several benchmarks. This directly challenges the common assumption that the best pruning is one consecutive chunk of middle-or-late blocks, and shows that reading off the ground state and low-lying excited states is essentially free, yielding a spectrum of high-quality candidates rather than one fragile answer. Uses a concrete excited-state example for Llama-3.1-8B-Instruct at 16/32 blocks removed and states that after retraining the configuration beats the ground state on several benchmarks; the source is an article-style overview, with details in the paper's Figure 2.

Perspective

The result targets engineering and research settings that need to deeply compress large language models under limited compute: build an approximate Hessian once on a small calibration dataset, then either brute-force on a single GPU or hand the QUBO to solvers such as open-source tabu search to generate candidate block-removal configurations, reusable across different compression targets M. It applies to homogeneous stacks and to hybrid architectures interleaving Mamba2, attention, and MoE, and is positioned as a pipeline component that composes with quantization, low-rank/SVD compression, width pruning, and knowledge-distillation-based healing.

The source is an article-style overview and does not give the specific size and composition of the calibration dataset, the error range of the Hessian approximation, complete score tables for each benchmark, the specific retraining setup, or the range over which solvers such as tabu and brute-force enumeration have been verified to agree on larger models; the boundary of the energy proxy being strong but not perfect, the stability of the excited-state advantage across models and compression rates, and the generality of uneven redundancy in hybrid architectures remain open questions worth confirming in the full paper.

Sources