PCD pins the primary gradient into the update direction and holds 70.9% accuracy at 90% sparsity on ResNet-34/CIFAR-100 while symmetric multi-objective baselines collapse
Synopsis
The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
Interpretation
PCD casts the multi-objective update as a small convex quadratic program: among all directions guaranteeing each secondary at least a τ-fraction of its maximum normalized first-order progress, it takes the one closest in Euclidean distance to the primary gradient, so the primary coefficient is structurally pinned to one and secondaries enter only through non-negative multipliers. Existing gradient-based multi-objective methods (weighted sum, MGDA, PCGrad, CAGrad and others) treat objectives as symmetric and interchangeable and target Pareto stationarity; PCD embeds the hierarchy directly in the update geometry, and its normalized direction vanishes only at composite multi-stationary (CMS) points, which are strictly stronger than Pareto stationarity. The paper provides a projection and KKT characterization, closed-form solutions for two objectives and via Cramer's rule for three, a proof of asymptotic scale invariance under EMA normalization, and a CMS endpoint theorem; the authors state that the theorem is an endpoint characterization of the deterministic normalized QP rather than a stochastic convergence claim.
In structured pruning, PCD sits above every general multi-objective baseline across eight (architecture, dataset) configurations at every compression target; on ResNet-34/CIFAR-100 it retains 70.9% accuracy at 90% sparsity (unpruned 73.3%), while MGDA falls to 1.2%, FAMO to 2.7%, PCGrad to 7.4% and CAGrad to 7.3%. The paper attributes this collapse to objective scale rather than gradient conflict: raw regularizer gradients carry roughly one to two orders of magnitude the norm of the task gradient, while the cosine between them never exceeds a small value; PCD's per-objective normalization keeps a single dimensionless τ interpretable across heterogeneous secondaries. Five seeds per configuration, 300 epochs, Adam with cosine annealing, means reported; Tables 1 and 2 give full numbers for DenseNet-121/CIFAR-10 and ResNet-34/CIFAR-100 including FLOPs, latency and file size.
Against tuned pruning-specific and constrained baselines under one identical protocol, PCD is better in 28 of 47 matched comparisons in the compression regime (above roughly 80% sparsity), worse in 1 and within seed noise in 18, with 25 of 47 baseline points Pareto-dominated by some point of PCD's own sweep. The comparison isolates training-time native sparsity as the only variable, excluding fine-tuning, per-layer scaling and stochastic gates, and pairs each baseline point with the PCD point of nearest mean sparsity, with accuracy never consulted in the matching. ResNet-34 and DenseNet-121 on CIFAR-10/CIFAR-100, 3 seeds, 300 epochs, 28 baseline operating points per configuration (330 runs); the paper notes the single loss is a genuine crossing where tuned scalarization reaches a given accuracy at a given sparsity on ResNet-34/CIFAR-100 while PCD is slightly lower at matched sparsity.
In a three-objective setting (cross-entropy plus L1 plus nuclear norm), PCD leads on the accuracy and effective-rank axes and starts less sparse on the parameter-sparsity axis before climbing to the baseline cluster; on ResNet-34/CIFAR-100 it holds 71.0% accuracy, 92.6% sparsity and effective rank 24.5, while baselines collapse in accuracy and over-compress. The paper stresses that the two secondaries are not aligned: L1 acts entry-wise while the nuclear norm acts on the spectrum and collapses entire singular directions; a single τ is meaningful because the constraint is stated on EMA-normalized gradients, demanding the same fraction of each secondary's own maximal normalized progress. ResNet-34 and an Inception-style model on CIFAR-10/CIFAR-100, 3 seeds, 300 epochs; beyond per-axis curves the paper reports dominated hypervolume and repeats the comparison under equal cardinality, one point per method, and alternative normalization bounds, with PCD leading under every variant.
Perspective
The result targets hierarchical optimization where secondary objectives are proper, convex and lower semi-continuous, typically group Lasso for structured pruning, unstructured L1 and nuclear-norm regularization; the authors position the framework for the regularizer-style hierarchy of Eq. (1), where secondaries are jointly stationary at primary minima. For practitioners, τ acts as a continuous compression dial at training time, moving from light to aggressive compression on ResNet-34/CIFAR-100, with the paper reporting the largest resource gains up to about τ = 0.1 and curves plateauing beyond that. The theoretical guarantees attach to the normalized direction and the idealized first-order step, and the authors state that the deployed update rescales to the raw primary-gradient magnitude for compatibility with standard optimizer step-size scheduling.
The authors state that they prove no convergence theorem for Algorithm 1, neither convergence of iterates to CMS points nor rates; the deployed update is zeroed by the rescale factor when the raw primary gradient vanishes, so primary-stationary but non-CMS fixed points exist, and the authors report this branch is never approached across all runs of the tuned-baseline study. Theorem 4.8 is a pointwise endpoint characterization and not a stability result, and its converse identification relies on automatic differentiation returning zero at subdifferentials, which the nuclear norm at the zero matrix does not satisfy. At the optimizer level, the paper reports that under Adam the per-step secondary-constraint sign survives on roughly 98.5%–99.8% of steps, while under SGD with momentum it flips on 47%–73% of steps, though both land on the same frontier, with the SGD control plateauing 20–40 epochs later. In addition, the QP can be infeasible when secondary constraints are strictly mutually incompatible; the authors report synthetic evidence that this is improbable in practice but still list it as a possible limitation, and multi-level priority hierarchies with adaptive τ schedules are left as future directions.
