Just Leaf It: Accelerating Diffusion Classifiers with Hierarchical Class Pruning
Synopsis
The work proposes a Hierarchical Diffusion Classifier (HDC) that exploits parent-child label hierarchies: in a pruning stage it traverses the label tree level by level, estimates epsilon-prediction errors with fewer Monte Carlo samples, and keeps only the lowest-error nodes, then runs classical diffusion classification on the surviving leaf nodes, achieving roughly 60% inference speed-up on ImageNet-1K with Stable Diffusion 2.0 (1600s to 650s) or raising per-class accuracy from 64.90% to 65.16% at nearly unchanged runtime.
Figure 2. Overview of our Hierarchical Diffusion Classifier (HDC). Starting with an input image x, noise ε ∼N(0, I) is added to generate a noisy image, resulting in xt for multiple timesteps t. Next, we use the diffusion classifier with a reduced number of ε-predictions and hierarchical conditioning prompts like “A photo of a {synclass / class name}” to progressively refine the classification through multiple levels of the label tree. By doing so, we keep track of the most promising classes (highlighted in green) and ignore the rest (highlighted in red). The set of selected nodes during the pruning stage is denoted as Sd
· Page 4Interpretation
Introduces a training-free Hierarchical Diffusion Classifier (HDC) that replaces evaluating all classes with hierarchical pruning followed by classification. Prior diffusion classifiers (e.g., Li et al., Clark et al.) evaluate every class per image, so inference scales linearly with the number of classes; HDC instead estimates epsilon-prediction errors for synsets in a pruning stage using fewer Monte Carlo samples, keeps top-k nodes according to a pruning ratio Kd, and runs classical diffusion classification only on the remaining leaf nodes. The paper provides the full procedure in Algorithm 1 and formal definitions in Equations (7), (8), and (9), with experiments on ImageNet-1K (a 7-level tree traversed from level 3) and CIFAR-100 (a self-generated tree).
On ImageNet-1K, the fixed pruning strategy (Strategy 1, Kd=0.5) raises per-class accuracy from 64.90% to 65.16% while cutting inference time by about 38.75%. Against the baseline diffusion classifier at 1600s and 64.90%, HDC Strategy 1 reaches 65.16% in 980s, i.e., nearly 40% less inference time without sacrificing accuracy, and is reported as new state-of-the-art accuracy for diffusion classifiers. Table 5 reports per-class average accuracy and time, and Table 1 reports Top-1/Top-3/Top-5 with time; evaluations use Stable Diffusion 2.0 at 512x512, l2 norm for epsilon-predictions, and timesteps uniformly sampled from [1, 1000].
The dynamic pruning strategy (Strategy 2, keeping nodes within two standard deviations of the minimum error) further reduces inference time to 650s, about 60% speed-up, with accuracy dropping to 63.33%. Compared with fixed pruning, dynamic pruning adapts candidate selection to the error distribution, offering a more speed-oriented trade-off and showing that HDC can be tuned continuously between speed and accuracy. Tables 1 and 5 report 650s, 59.38% speed-up, and Top-1 of 63.20% (Table 1) and per-class 63.33% (Table 5); on CIFAR-100, Strategy 2 (Kd=0.4) reports +3.3 percentage points accuracy and about 34% speed-up.
Pruning ratio, diffusion model version, and prompt template all affect the speed-accuracy trade-off, and the default prompt 'A photo of a' performs best. The paper compares SD 1.4/2.0/2.1 and several prompt templates, showing SD 2.0 under Strategy 1 attains the highest Top-1 (64.14% at 980s), while templates such as 'A bad photo of a' and 'A low-resolution photo of a' reduce accuracy. Table 2 reports per-class/overall Top-1 and time for three SD versions under both strategies; Table 4 reports Top-1/Top-3/Top-5 for four prompt templates under both strategies.
Perspective
The result applies to datasets with well-defined hierarchical labels (such as the WordNet-based ImageNet-1K, where the paper uses a 7-level tree traversed from level 3) and to CIFAR-100 with a self-generated hierarchy; the method is training-free, can be layered onto existing conditional diffusion models (the paper validates SD 1.4/2.0/2.1), and supports adding or removing class labels without retraining. For readers who want to control inference cost in large-scale classification or to tune the speed-accuracy balance per scenario, HDC offers an actionable route.
The paper notes that efficiency gains depend on the depth and balance of the label tree, so datasets with shallow hierarchies or weak parent-child relationships may benefit less; hierarchies could be built bottom-up with LLMs or refined via greedy expansion, but the paper does not provide a full evaluation of these alternative constructions. Inference time also varies markedly across classes (e.g., 'snail' at 221s versus 'keyboard space bar' at 1400s), and prompt templates and SD versions shift the trade-off, so readers transferring the method to their own data should verify the effects of pruning ratio and hierarchy quality.
