SRMA-Mamba lifts cirrhotic liver MRI segmentation to 92.95% mDSC on T1W and 86.25% on T2W in CirrMRI600+ while cutting compute
Synopsis
The work introduces SRMA-Mamba, a Mamba-based network that uses the Spatial Anatomy-Based Mamba (SABMamba) module to perform selective scans across sagittal, coronal, and axial anatomical planes for a global spatial context, and the Spatial Reverse Mamba Attention (SRMA) module to progressively refine boundaries from a coarse segmentation map and hierarchical encoder features, reporting better segmentation metrics than SegResNet, UX-Net, MedNeXt, SwinUNETR, SwinUNETRv2, and SegMamba on CirrMRI600+ T1W and T2W, with fewer parameters and GFLOPs than SegMamba.
Interpretation
It proposes SRMA-Mamba, a Mamba-based encoder-decoder network for volumetric pathological liver segmentation in MRI, where the encoder extracts hierarchical features and the decoder progressively refines the segmentation map using the coarse map and features from each encoding stage. Earlier Mamba-based medical segmentation methods such as SegMamba treat the whole volume as a linear sequence; the paper argues this linearization disrupts the intrinsic 3D spatial topology, which matters because cirrhotic structures are highly irregular and anisotropic. The paper presents the overall architecture and the encoder-decoder pipeline, and compares against CNN, Transformer, and Mamba methods on CirrMRI600+ T1W/T2W using the same codebase, dataset partitioning, and hyperparameter tuning.
It proposes the SABMamba module, which performs selective Mamba scans across the three orthogonal anatomical planes and fuses plane information into a global spatial context representation. It translates the radiologist habit of reading MRI volumes from multiple anatomical planes into multi-plane scanning inside the network, rather than a single linear sequence scan. Ablation shows that replacing the SABMamba encoder with a UNETR encoder drops mDSC from 92.95% to 80.62%; in the plane-configuration comparison the three-plane setting (mDSC 92.95%, mIoU 87.59%, HD95 17.52) outperforms axial (92.48%), sagittal (92.50%), and coronal (90.22%) single-plane variants.
It proposes the Anatomy-Based Selective Scan (ABSS) module, which directly processes 3D voxel data to produce 3D feature representations, using the S6 selective mechanism to filter irrelevant information while preserving relevant features. It processes volumetric data directly at low computational complexity while preserving spatial anatomical information, instead of collapsing the volume into a linear sequence. In the complexity table SRMA-Mamba uses 17.22M parameters and 149.14 GFLOPs, below SegMamba's 64.24M and 2379.03 GFLOPs, and below the GFLOPs of SwinUNETR, SwinUNETRv2, MedNeXt, and UX-Net.
It proposes the Spatial Reverse Mamba Attention (SRMA) module, which uses reverse attention weights Ai = 1 − Sigmoid(Si+1) together with encoder features to refine boundary details stage by stage. It combines reverse attention with Mamba for boundary refinement across multiple decoding stages, letting the coarse segmentation map and hierarchical features complement each other. Ablation shows that removing SRMA reduces mDSC from 92.95% to 90.63%; on T1W, HD95 is 17.52 and ASSD is 3.05, better than SegMamba's 26.97 and 4.69.
Perspective
The result targets segmentation of pathological liver structures in MRI volumes and applies to abdominal MRI settings similar to CirrMRI600+, a single-center, multi-vendor, multi-sequence dataset, for both T1W and T2W. It lets follow-up work build on the combination of multi-plane anatomical scanning and reverse-attention boundary refinement, and makes it easier for researchers to compare Mamba-based volumetric segmentation models on the same benchmark; the code is released at https://github.com/JunZengz/SRMA-Mamba for reproduction and further development.
The paper notes challenging cases in T2W that occasionally produce false positives and missed detections of cirrhotic tissue, indicating boundary determination remains difficult in that modality; the dataset is single-center, multi-vendor, and multi-sequence, with test sets of 31 T1W and 31 T2W cases, so behavior when generalizing to other centers, protocols, or populations still needs further observation. Qualitative results are presented as images, and case-level error distributions are not expanded in the main text.
