Skip to main content
Back to timeline
arXivSource publication:

FermatSyn combines SAM2 priors with Fermat-spiral Mamba scanning to synthesize missing medical modalities, topping four brain-imaging benchmarks while segmenters trained on its synthetic images match real-image training

Synopsis

The work proposes FermatSyn, which injects anatomical priors via LoRA+ fine-tuning of a frozen SAM2 vision encoder, preserves high-frequency lesion detail with HRDM and CIN, and builds an approximately isotropic receptive field through continuity-constrained Fermat spiral scanning inside a bidirectional Mamba; on SynthRAD2023 and merged BraTS (including BraTS-MEN and BraTS-MET) it surpasses compared methods on PSNR, SSIM, FID and 3D structural consistency, and segmentation models trained on its synthesized images show no significant difference from real-image training (p>0.05).

AI-generated editorial illustration: FermatSyn: SAM2-Enhanced Bidirectional Mamba with Isotropic Spiral Scanning for Multi-Modal Medical Image Synthesis

Interpretation

A SAM2-based prior encoder: the pretrained SAM2 vision transformer is frozen and only LoRA+ low-rank adapters update its MLP and MHSA layers, transferring structural representations such as organ boundaries, tissue interfaces and shape saliency learned for segmentation into the synthesis task. Prior medical image synthesis frameworks lacked a mechanism to inject domain-aware anatomical knowledge, which the authors associate with cross-modal implausibility; this work explicitly wires SAM-family segmentation priors into the synthesis pipeline. Ablation shows SAM2-VTE is the largest single gain, adding 15.5% PSNR, with cumulative improvement of about 21% over the GAN baseline (Table 5a).

An HRDM detail encoder plus a CIN cross-scale integration network: HRDM uses three parallel dilated convolutions (rates 1, 3, 5), a depthwise separable convolution and a high-pass filter to retain high-frequency texture, while CIN splits channels into even and odd groups processed by 5×5 and 3×3 depthwise convolutions for low-frequency semantics and local structure respectively. Addressing the loss of high-frequency detail needed for small lesions under conventional pooling and underdeveloped cross-scale fusion, it offers explicit high-frequency compensation and channel-split fusion. In the step-wise ablation, adding HRDM+CIN on top of SAM2-VTE raises SynthRAD PSNR from 27.35 to 27.74 and BraTS from 27.18 to 28.45 (Table 5a).

A continuity-constrained Fermat spiral scanning strategy with a bidirectional BFS-Mamba: 2D feature maps are serialized along a Fermat spiral with a golden-angle step of about 137.508°, a grid-matching objective (Eq. 6) trades off global isotropy against local path continuity, and forward and backward SSM paths are fused. The authors report that the rectangular spiral of I2I-Mamba still shows corner hot-spots (operator footprint σ=0.124), whereas the Fermat spiral lowers Delaunay variance from 0.0154 to 0.0061 (a 60% reduction) and reduces Jacobian σ by 29%. In an architecture-controlled scanning comparison, the Fermat spiral beats the rectangular spiral by 1.29/1.40 dB PSNR on SynthRAD/BraTS (Table 5b); λc=0.7 was best on the validation set among {0.3, 0.5, 0.7, 0.9}.

Validation across four benchmarks covers both synthesis quality and downstream clinical usability: brain tumor segmentation Dice is 0.847/0.762/0.785 (WT/ET/TC) for T1n→T2w and 0.851/0.775/0.798 for T1n→T2f, within 3% of the real-image topline, with no significant difference across all three tumor regions and both tasks (p>0.05). Most compared methods reach p>0.05 only for the structurally most complex TC region, whereas this work extends that equivalence to the harder WT and ET regions. Two protocols are used, real-trained/synthetic-evaluated and synthetic-trained/real-evaluated, with Wilcoxon signed-rank tests; inference is 31 ms per 256×256 slice with 87.3 M trainable parameters.

Perspective

The results target brain imaging settings where a missing modality must be filled in: inter-sequence brain tumor MRI translation (T1n/T2w/T2f→T1c) and MRI↔CT translation on SynthRAD2023, with data filtered for SNR<15 dB and motion artifacts above grade 3, the central 80 slices per volume retained, and 256×256 center cropping. For teams aiming to reduce acquisition burden in radiotherapy planning, surgical navigation or scarce-data research, this offers reproducible code and architecture; the authors also list extending the Fermat spiral to 3D volumetric synthesis and distilling SAM2-VTE into lightweight student networks for real-time deployment as next steps.

Readers should still watch: validation centers on brain MRI and MRI–CT, so whether the same holds for other anatomical sites and modality combinations remains to be tested; downstream segmentation uses only a ResNet50-based U-Net and the WT/ET/TC regions, leaving other segmentation backbones and finer-grained lesions open; the Fermat spiral currently operates on 2D slices and the authors list 3D volumetric synthesis as future work, so how inter-slice consistency gains change under true 3D serialization is an open question; and hyperparameters such as λc and the LoRA+ α were selected on the validation set, so sensitivity under cross-dataset transfer needs further confirmation.

Sources