ProSMA-UNet replaces U-Net skip gating with an ℓ1 proximal sparse gate, reporting best results across 2D and 3D medical segmentation benchmarks and about a 19% F1 gain on 3D colon.
Synopsis
The work proposes ProSMA-UNet, which recasts skip connections in U-shaped medical segmentation networks as a decoder-conditioned sparse feature selection problem: it builds a multi-scale encoder-decoder compatibility field with lightweight depthwise dilated convolutions, applies an ℓ1 proximal operator with learnable per-channel thresholds to yield a closed-form soft-thresholding gate, and adds decoder-conditioned channel gating driven by global decoder context, reporting best results on three 2D benchmarks (BUSI, GlaS, Kvasir-SEG) and two 3D benchmarks (Spleen and Colon from the Medical Segmentation Decathlon), with roughly a 19% relative F1 gain over the strongest baseline on Colon.
Fig. 1. ProSMA-UNet motivation and overview. (a) Conventional attention gates generate a dense soft mask (sigmoid reweighting), which can still pass weak but harmful skip activations (noise features) into the decoder. (b) Our ProSMA sparse gating constructs a multi-scale decoder–encoder compatibility field and applies an ℓ1 proximal (soft-thresholding) operator to induce explicit sparsity (exact zeros), enabling direct removal of irrelevant skip responses. (c) ProSMA-UNet integrates the proposed sparse gating module at each skip connection to condition high-resolution feature transfer on decoder context. (d) Overview of ProSMA Spare Gating.
· Page 4Interpretation
The paper formalizes skip gating as a decoder-conditioned sparse feature selection problem rather than dense reweighting, and identifies skip connections as a major pathway for noise and background leakage into the decoder. In contrast to sigmoid dense masks such as Attention U-Net, the work distinguishes attenuation from removal and treats the skip pathway as a selection operator Gs:(xs,gs+1)->x̃s that retains only components relevant under decoder context. This is a problem reformulation and motivation argument, supported by the schematic contrast in Figure 1(a)(b) showing that an attention mask cannot filter noise while the proximal sparse gate can zero it; it is conceptual rather than quantitative.
The ProSMA gate consists of a multi-scale compatibility field followed by ℓ1 proximal sparsification, yielding a closed-form soft-thresholding rule that sets incompatible activations to exact zeros. The compatibility field u=fms(ReLU(q+k)) aggregates multiple receptive fields via depthwise dilated convolutions while preserving channel independence to match per-channel sparsity; sparsity comes from z*=prox_{λ‖·‖1}(u)=sign(u)max(|u|−λ,0), with λ=softplus(θ) as a learnable per-channel threshold. The method is fully derived with Equations (1)-(5) and a closed-form solution; sparsity follows directly from the dead-zone property of soft thresholding, making this an analytical rather than empirical result.
Theorem 1 proves that proximal sparse gating achieves exact feature selection and stability: increasing the threshold can only remove active features, and the operator is non-expansive (1-Lipschitz), so it cannot amplify noise in the compatibility field. Compared with attention gates validated mainly empirically, this work provides formal guarantees of monotonic sparsity and non-expansiveness, with a proof via entrywise soft thresholding and a Frobenius-norm comparison. The theorem and proof appear in the main text, relying on the piecewise-linear soft-threshold function with slopes in [0,1]; this is a theoretical analysis.
The method reports consistently best results on three 2D and two 3D medical segmentation benchmarks, with the largest gains on difficult 3D tasks; ablations show spatial sparsity and channel gating are complementary. On BUSI it improves over the strongest competitor by +2.86 IoU / +2.43 F1; on GlaS it reaches 88.76±0.14 IoU / 94.04±0.08 F1; on Kvasir-SEG it reaches 66.23±2.15 IoU / 85.43±1.19 F1; in 3D it reports 97.59±0.10 on Spleen and 63.14±0.99 on Colon, the latter +10.09 (about +19.0% relative) over UKAN2.0 3D. In ablation, the full model reaches F1 80.13, +6.08 over the ungated 74.05. 2D results are reported over three independent runs with standard deviations; 3D results are reported as F1; the ablation compares four gating configurations in a single setting. Some baseline results are taken from U-KAN under the same protocol, and UKAN2.0 could not be evaluated on GlaS due to GPU memory limits.
Perspective
The result targets medical image segmentation with U-shaped encoder-decoder backbones, in low-contrast and noisy 2D and 3D clinical imaging settings such as breast ultrasound, histology, endoscopic polyps, and abdominal CT spleen and colon tumor segmentation. For researchers and practitioners who want to reduce skip-connection noise leakage without changing the overall backbone, the ProSMA gate can serve as a drop-in module at skip connections; the paper also provides a closed-form soft-thresholding implementation and theorem guarantees, making it straightforward to embed in existing pipelines.
The paper reports performance on the listed benchmarks; behavior on other anatomical targets, imaging modalities, and cross-center data remains to be observed. 3D results are reported as F1 while 2D results are reported as means and standard deviations over three runs, so evaluation conventions differ across tasks. The ablation compares four gating configurations in a single setting, and the distribution of the sparsity threshold λ and how the learnable thresholds evolve during training are not expanded in the main text. Some baseline results follow the same protocol as U-KAN, and UKAN2.0 could not be evaluated on GlaS due to GPU memory limits, so the comparison scope on that dataset is limited. In addition, this reading is of the full text, but the qualitative visualizations and error maps in Figure 2 are described only in prose, so case-level differences cannot be verified item by item from the text.
