Skip to main content
Back to timeline
arXivSource publication:

Letting the Sparsity Coefficient Learn Its Own Compression: Training-Adaptive Convolutional Sparse Coding Through an Information Bottleneck Lens

Synopsis

This work turns the sparsity coefficient of convolutional sparse coding from a manually fixed hyperparameter into a differentiable variable jointly learned with network parameters through unfolded FISTA iterations, interprets that coefficient within an information bottleneck framework as the controller of the compression-retention trade-off, achieves competitive clean-data recognition on CIFAR-10/100 and ImageNet-1K with greatly improved robustness under various input perturbations, and introduces a label-free post-training adaptation strategy that adjusts compression strength using a small set of unlabeled corrupted samples while keeping the main network parameters frozen.

AI-generated editorial illustration: Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

Interpretation

It proposes training-adaptive convolutional sparse coding (TA-CSC), embedding the sparsity coefficient as a differentiable variable inside unfolded FISTA iterations and optimizing it jointly with the convolutional dictionary and network parameters. Earlier CSC-style methods such as ML-CSC, SCN and SDNet typically treat the sparsity coefficient as a pre-selected hyperparameter; here it becomes a learnable quantity determined jointly by the representation and the downstream task, with a complete recursive differentiation path through the unrolled iterations. The paper provides per-iteration gradient derivations in Eqs. (11)-(14) and enforces non-negativity through parameterization; experiments build TA-CSC-18 and TA-CSC-18all on ResNet-18, use two unrolled FISTA iterations, and set the compression incentive to 0.001.

It establishes an explicit correspondence between CSC and the information bottleneck objective, treating the sparsity term as compression and the reconstruction term plus task loss as preservation of task-relevant information, with their competition implementing the compression-retention trade-off. Whereas CSC has mostly been used as a signal-processing or differentiable-optimization layer, this work explicitly reads the sparsity coefficient as the IB compression control variable and designs the training objective in Eqs. (15)-(16) accordingly. The correspondence is supported at the formulation level and by training-dynamics visualization: coefficients stay small early when task fitting dominates, then rise rapidly and converge as accuracy saturates, showing a fitting-to-compression transition; layer-wise, earlier layers show smaller coefficients and layers closer to the prediction objective show larger ones.

It introduces a label-free post-training adaptation strategy that freezes the main network parameters and adjusts only the compression coefficients using a small set of unlabeled corrupted or shifted samples, with relative reconstruction error as a fidelity measure. Unlike learning a fixed compression level once on clean data, this method re-estimates compression strength when the input distribution is corrupted, requires no labels, and differs from per-sample tuning of a fixed sparse-coding model. It uses 100 unlabeled corrupted samples by default; on CIFAR-10-C and ImageNet-C under Gaussian, speckle and impulse noise, TA-CSC-18 with post-training adaptation surpasses SDNet-18 with per-sample tuning, for example 66.63% versus 64.92% under CIFAR-10-C Gaussian noise.

It achieves competitive clean-data recognition and robustness on corrupted data simultaneously, and observes that the compression coefficient increases monotonically with corruption severity. Under the same training protocol, TA-CSC-18all reaches 97.65%, 80.76% and 72.53% on CIFAR-10, CIFAR-100 and ImageNet-1K, above ResNet-18 and CSC baselines such as SCN and SDNet; on robustness, TA-CSC-18 already consistently outperforms baselines even without post-training adaptation. Results come from classification experiments on CIFAR-10/100 and ImageNet-1K plus multiple noise settings on CIFAR-10-C and ImageNet-C; ablations show that increasing FISTA iterations from 2 to 8 and adaptation samples from 50 to 500 both help but with relatively limited gains, so 2 iterations and 100 samples are used by default.

Perspective

The results target image classification with a ResNet backbone, validated on CIFAR-10/100 and ImageNet-1K and their corrupted versions; the method replaces convolutional layers with FISTA-unrolled layers, and post-training adaptation suits settings where the input distribution is corrupted or shifted and a small number of unlabeled samples is available. For practitioners who want to improve robustness on corrupted data without retraining the backbone, and for researchers interested in controllable representation compression, this framework offers a directly reusable interface.

The paper itself notes that evaluation is mainly on the ResNet architecture and the classification task, and that the relationship between the learned coefficient and information compression rests primarily on empirical evidence, leaving a more rigorous theoretical link between sparse coding and the information bottleneck objective to be established. In addition, although the homepage evidence bundle is a full-text parse, figures appear here as text and tables, so the exact curve shapes of Fig. 2 and Fig. 3 can only be understood from the prose; behavior under other backbones, more visual tasks and broader distribution shifts, and whether the four-peak compression pattern during training is general, remain open questions worth watching.

Sources