Skip to main content
Back to timeline
arXivSource publication:

MAGEFormer embeds voxel spacing into positional encoding and attention, reaching the lowest HD95 on anisotropic abdominal CT segmentation across BTCV and FLARE 22

Synopsis

The authors propose MAGEFormer, which embeds physical spacing constraints into representation learning: MASE calibrates positional frequencies by voxel spacing, GCA suppresses physically implausible cross-scale feature correlations, and GVV performs deterministic sub-voxel view aggregation at inference; under a unified 5-fold protocol on BTCV and FLARE 22 it attains the strongest boundary accuracy among compared methods (10.58 mm and 3.40 mm HD95) with consistent Dice gains, and its advantage grows with anisotropy severity.

Source-provided article image: MAGEFormer: Learning Metric-Consistent Representations for Anisotropic CT Segmentation
Figure 1 ·

Figure 1: MAGEFormer: Architecture Overview and Geometric Motivation

arXiv

Interpretation

MAGEFormer attains the lowest HD95, i.e. the strongest boundary accuracy, among the compared methods on both BTCV and FLARE 22, with consistent Dice gains: 93.91% Dice / 3.40 mm HD95 on FLARE 22 and 84.11% / 10.58 mm on BTCV. Prior ViT-based volumetric segmentation largely relies on positional encodings and relative position biases parameterized on voxel indices, or on thickness-aware operator design, without enforcing metric consistency at the representation level; this work applies physical spacing constraints jointly at positional encoding, cross-scale feature interaction, and inference-time aggregation. Compared against nnU-Net, TransUNet, UNETR, Swin-UNETR, and U-Mamba under a fixed 5-fold cross-validation with a shared fold split, official model and parameter configurations, and a unified metric pipeline; cross-fold standard deviations are reported, including Dice std of 0.20 and HD95 std of 0.48 on FLARE 22.

Ablation shows each component helps and they compose: with Swin-UNETR as baseline, adding MASE, GCA, or GVV alone yields roughly +4.91%, +5.06%, and +3.84% average Dice respectively, MASE+GCA gives +8.08%, and the full model gives +9.24% (0.7487 to 0.8411), with HD95 dropping from 75.21 mm to 10.58 mm. The work decomposes metric consistency into three orthogonal levels—encoding, feature interaction, and inference—and validates them separately, indicating that geometric calibration is a design principle applicable at multiple entry points of geometric information rather than a single-point fix. A per-component ablation table on BTCV with Swin-UNETR as baseline, reporting corresponding HD95 changes, where MASE and GCA individually already produce marked HD95 reductions.

Gains concentrate on small and topology-sensitive organs vulnerable to anisotropic sampling: on BTCV, adrenal gland mean Dice reaches 72.76% and pancreas improves to 82.59%, the best among the compared methods. These structures are highly sensitive to through-plane resolution and partial-volume effects, so the results support that metric-consistent modeling improves fine-grained anatomical delineation rather than only aggregate overlap metrics. A per-organ Dice comparison table covering 13 abdominal organs; the authors note that boundary quality of low-contrast structures such as adrenal glands and pancreas tail is highly sensitive to sampling geometry.

The advantage widens as anisotropy becomes more severe: after stratifying by natural anisotropy ratio, the gain over baseline broadens as the ratio increases, and under synthetic z-degradation applied at test time only, MAGEFormer degrades more gracefully than compared baselines; GVV matches standard TTA accuracy at substantially lower inference time. This directly tests the central hypothesis that internal geometry calibration is more effective than relying on conventional isotropic preprocessing alone, and quantifies the efficiency advantage of deterministic aggregation over random TTA at inference. Analysis pooling test cases from both benchmarks stratified by anisotropy ratio, plus a controlled test-time perturbation that downsamples along z and interpolates back to the original grid, with shaded regions denoting cross-fold standard deviation; the speed-accuracy comparison comes from GVV versus standard TTA.

Perspective

The results target abdominal multi-organ CT segmentation under acquisition conditions where through-plane spacing is markedly coarser than in-plane spacing, validated on two benchmarks, BTCV (30 scans) and FLARE 22 (50 labeled scans), under a unified 5-fold protocol. The method is designed to remain compatible with standard U-shaped Transformer backbones, so it can be embedded into existing volumetric segmentation pipelines; the authors state that future work will extend this geometry-aware paradigm to other inherently anisotropic modalities such as multi-parametric MRI. For a reader, this means that when acquisition itself is clearly anisotropic and small organs or thin-structure boundaries matter, writing physical spacing into positional encoding, skip interaction, and inference aggregation is a reusable engineering path.

Readers should still watch: the anisotropy stratification and synthetic z-degradation experiments present stratification thresholds and curve values as figures, and the text does not give a complete numeric table; the number of GVV views and the per-view displacement are denoted symbolically in the text, so exact values require the original; the initial learning rate and input crop size among training hyperparameters are not shown numerically in the text; and the conclusions rest on two abdominal CT benchmarks, so generalization to other anatomical sites, other anisotropic modalities, and different acquisition protocols still needs independent validation.

Sources