Skip to main content
Back to timeline
arXivSource publication:

TranSamba uses cross-plane Mamba to push weakly supervised volumetric segmentation to 64.1% DSC, topping BraTS, KiTS and LASC

Synopsis

The work proposes TranSamba, which places a Cross-Plane Mamba (CPM) block before the Transformer block in every layer of a DeiT-S Vision Transformer encoder so that patch tokens in neighboring planes exchange information in linear time, improving in-plane self-attention and the class-to-patch attention maps; trained with only plane-level labels, 16 contiguous planes per volume and 16 volumes per batch for 100 epochs, it reaches 64.1%, 26.8% and 41.0% DSC and 46.8%, 25.8% and 25.6% IoU on the BraTS, KiTS and LASC test sets, the highest among the compared methods, with time complexity linear in the number of planes and batch space complexity independent of it.

Source-provided article image: Hybrid Transformer-Mamba for Weakly Supervised Volumetric Medical Segmentation

Interpretation

It proposes TranSamba, a hybrid architecture in which a CPM block and a Transformer block are arranged in series in every layer, with CPM handling cross-plane modeling and the Transformer handling in-plane self-attention and producing the C2P attention maps. Existing weakly supervised volumetric segmentation methods often rely on 2D encoders and neglect the volumetric nature of the data; cross-plane self-attention would model that volume but scales quadratically with the number of planes N, whereas Mamba offers a linear-time alternative. The paper derives the complexity: TranSamba's time complexity is 128MND+4MD²+2M²D, linear in N, while replacing the cross-plane SSM with SA gives 4MND²+2M²N²D+4MD²+2M²D, quadratic in N; space complexity B(4MD²+2M²D)+(B/N)(128MND) is independent of N.

Cross-plane modeling, not added model complexity, is the main source of the improvement. The authors introduce V2 as a control with complexity comparable to V3 (alternating Transformer and in-plane Mamba blocks, equivalent to the forward branch of Vision Mamba), separating the effect of complexity from that of cross-plane modeling. In the ablation, V3 improves substantially over the ViT-only baseline V1 (BraTS 60.7% vs 40.7% DSC, KiTS 25.1% vs 15.8%, LASC 45.9% vs 33.5%), whereas the comparable-complexity V2 improves far less (41.9%/19.5%/38.1%); the Mamba-only variants V4 and V5 perform substantially worse, though V5 still beats V4.

Across three datasets covering different modalities and pathologies, TranSamba achieves higher DSC and IoU than the compared weakly supervised methods. AME-CAM and UM-CAM perform dynamic multi-scale feature aggregation and IAT uses focal loss plus self-supervision; these strategies are effective on BraTS but drop substantially on KiTS and LASC, whereas TranSamba is more consistent across the three. On the test set TranSamba reaches BraTS 64.1/46.8, KiTS 26.8/25.8 and LASC 41.0/25.6 (DSC/IoU), exceeding the second-highest performer of each dataset by mean margins of 5.3% (DSC) and 2.4% (IoU); compared methods include Grad-CAM, Grad-CAM++, FullGrad, Ablation-CAM, Score-CAM, SEAM, LayerCAM, TS-CAM, SIPE, AME-CAM, UM-CAM and IAT.

Plane count and layer order matter: local cross-plane modeling over 16 planes performs best, and the CrossIn design attains the tied-highest ranking of 1.83. The authors systematically compare 16/32/64/128 planes and the Parallel, InCross and CrossIn layer designs, and report HD95 to separate variants with close DSC/IoU. 16 planes gives the highest DSC on all three datasets (BraTS 60.7, KiTS 25.1, LASC 45.9); CrossIn's HD95 is 4.6/15.2/3.2 for BraTS/KiTS/LASC, versus Parallel 4.8/15.5/3.1 and InCross 3.8/15.4/3.6.

Perspective

The result targets binary volumetric segmentation trained from plane-level labels, for continuous structures spanning a limited number of planes in CT and MRI, such as the whole tumor on BraTS T2 FLAIR, the kidney tumor in KiTS 2023 and the LA cavity in LASC; the training configuration is 16 volumes per batch with 16 contiguous planes each, 100 epochs, on a single 40GB A100. It lets follow-up work obtain better localization maps through architecture alone, without multi-scale aggregation, boundary/size handling, self-supervised regularization or target-specific constraints; the authors note future extension to multi-class scenarios and the addition of auxiliary techniques for more clinically oriented applications.

Readers should still watch: the task is limited to binary segmentation, while multi-class segmentation needs orthogonal techniques such as co-occurrence or inter-class or inter-channel correlations; performance correlates positively with object size, with KiTS tiny and small subgroups at only 5.1% and 21.3% DSC, so small objects remain an open question; on the KiTS huge subgroup FullGrad's 66.1% exceeds TranSamba's; LASC has only 43 training studies and the KiTS and LASC backgrounds are noisy cavities; the authors state that TranSamba still requires post-processing or multi-step approaches for clinical applications, and that the lower speed of V4 and V5 stems from implementation inefficiency. In addition, this is a fast parse: the specific content of Figures 1 and 2 and some table layout details cannot be fully reconstructed from the text, so the qualitative comparison and the size-DSC correlation should be checked against the original figures.

Sources