Skip to main content
Back to timeline
arXivSource publication:

BiM-GeoAttn-Net reaches 93.35% Dice on 71 aortic dissection CTA cases, beating CNN, Transformer, and SSM baselines

Synopsis

The work proposes BiM-GeoAttn-Net, which cascades a Bidirectional Depth Mamba (BiM) and a Geometry-Aware Vessel Attention (GeoAttn) module at the nnU-Net bottleneck, and on 71 multi-source Stanford Type-B aortic dissection CTA cases performs binary segmentation of the vessel foreground (true plus false lumen), achieving Dice 93.35%, IoU 87.53%, Recall 94.85%, Precision 92.04%, and HD95 12.36 mm, outperforming Attention U-Net, nnU-Net, Swin-UNet, SegFormer3D, and Mamba-UNet on overlap metrics while keeping boundary accuracy and computational cost competitive.

Source-provided article image: BiM-GeoAttn-Net: Linear-Time Depth Modeling with Geometry-Aware Attention for 3D Aortic Dissection CTA Segmentation
Figure 1

Figure 1: Challenges in AD CTA segmentation. Four consecutive axial slices exhibit strong cross-slice continuity, subtle boundary contrast, and morphological variation. Ground-truth masks are in red.

· Page 2

Interpretation

BiM-GeoAttn-Net is a lightweight 3D segmentation framework that adds two complementary modules at the nnU-Net bottleneck to target limited long-range context modeling and low-contrast boundary ambiguity. Unlike stacking full 3D self-attention or generic sequence modeling, the design places depth-axis state-space modeling and vessel geometry priors in the same cascaded bottleneck stage. The paper provides a full architecture diagram and formal module descriptions, and compares against five representative baselines under identical training settings.

The Bidirectional Depth Mamba (BiM) performs bidirectional state-space scanning along the depth axis, aggregating cross-slice context at near-linear complexity to stabilize inter-slice continuity. Existing SSM-based medical segmentation methods typically use generic sequence construction without aligning the modeling direction with anatomical continuity; BiM fixes the sequence direction to the depth axis and fuses both directions. Ablation shows adding BiM alone raises Dice from 89.77% to 91.12% and lowers HD95 from 29.08 mm to 20.35 mm.

Geometry-Aware Vessel Attention (GeoAttn) uses three plane-aligned anisotropic convolutions plus a standard 3D convolution to capture orientation-sensitive tubular structure, then refines boundaries via spatial and channel attention. Compared with purely global modeling, the module explicitly injects vessel-structure priors to sharpen ambiguous boundaries under low contrast. Adding GeoAttn alone gives Dice 90.67% with improved Recall and HD95 25.48 mm; combining both modules yields Dice 93.35% and HD95 12.36 mm.

On 71 multi-source TBAD CTA cases split at the patient level, the method leads across overlap metrics while training cost stays close to nnU-Net. Relative to nnU-Net, Dice and IoU improve by 2.51% and 3.89%; relative to Transformer-based models, training time and memory are lower. The test set has 14 cases, with 50 training and 7 validation cases; 2.9 min/epoch and 8.0 GB memory versus nnU-Net's 2.8 min and 7.6 GB.

Perspective

The result targets binary segmentation of the vessel foreground (true plus false lumen) in Stanford Type-B aortic dissection CTA, using data from two public sources, ImageTBAD and TBD-CTA, split at the patient level into 50/7/14 cases. The method uses an nnU-Net 3d_fullres backbone trained with 128×128×128 patches, batch size 2, 400 epochs, and SGD with cosine annealing, suited to settings that must handle anisotropic 3D CTA under limited GPU memory while caring about inter-slice continuity and tubular boundaries. For researchers wishing to reuse the idea, BiM and GeoAttn can serve as bottleneck enhancement modules plugged into other U-shaped 3D networks; for clinical or engineering readers, the value lies in higher overlap accuracy at roughly nnU-Net cost.

The test set has only 14 cases, and HD95 is not the best among all methods (Attention U-Net reaches 10.77 mm), so the relative boundary advantage needs observation on larger, multi-center data. The paper excludes false-lumen thrombosis annotations as an independent class and performs only binary vessel-foreground segmentation, so discrimination of thrombotic regions remains unclear. The paper also reports no statistical significance tests or confidence intervals; cross-source generalization, stability under different scanning protocols, and whether downstream morphological quantification or clinical decisions actually improve remain open questions.

Sources