DA-SAM3 routes over dual-adaptive low-rank experts to gain about 5% accuracy while cutting MoE parameter overhead by over 80% on four public medical segmentation benchmarks
Synopsis
The work proposes Dual-Adaptive SAM3 (DA-SAM3), which replaces selected feed-forward blocks in SAM3's fusion module with dual-adaptive MoE layers: a task-aware Dynamic Expert Router (DER) sparsely activates experts by jointly reasoning over visual content and the textual concept prompt, while a parameter-aware Decomposed Parameterized Experts (DPE) design represents each expert as a shared frozen base inherited from pretrained SAM3 plus a lightweight trainable low-rank delta, so that on the four public datasets Synapse CT, MMWHS, BTCV and ACDC it matches or exceeds fully fine-tuned SAM3 and standard MoE baselines, reports roughly a 5% gain over then state-of-the-art methods, and reduces MoE parameter overhead by more than 80%.
Fig. 1. (a) The architecture of the Dual-Adaptive SAM3 framework. (b) The architec- ture of trainable Transformer. (c) The Dual-Adaptive MoE Layer including Dynamic Expert Router (DER) and Decomposed Parameterized Expert (DPE)
· Page 3Interpretation
It introduces DA-SAM3, a framework that combines dynamic multimodal routing with parameter-efficient expert decomposition to adapt a concept-conditioned vision-language foundation model to medical image segmentation. The authors describe it as the first framework to integrate dynamic multimodal routing with parameter-efficient expert decomposition for adapting vision-language foundation models to medical imaging, in contrast to static and uniformly applied LoRA/Adapters and to costlier conventional MoE. The paper provides a full method description (DER, DPE, hierarchical placement, two-stage training) plus comparisons and ablations on four public datasets; the evidence is the authors' own reported tables and visualizations, without external reproduction.
The Dynamic Expert Router (DER) conditions on a global domain-context vector and local token features to sparsely select top-k experts per token (k=2), using StopGrad to protect the frozen concept embedding. Compared with random routing or simply concatenating visual and text tokens, DER explicitly uses clinical concept embeddings to activate task-relevant experts; in ablation, replacing DER with random routing clearly degrades performance. The ablation table shows random routing reaching DSC 82.75 and HD 12.120 on Synapse CT versus 85.12 and 8.425 for the full model; this supports the router's role within the authors' experimental setting.
Decomposed Parameterized Experts (DPE) writes each expert as a shared frozen pretrained FFN weight W0 plus a low-rank delta ΔW=AB^T, reducing MoE parameter overhead by more than 80%. Rather than implementing experts as full dense networks, DPE reuses pretrained weights as a shared base and trains only low-rank deltas, preserving expert specialization while sharply compressing added parameters. The abstract and introduction report over 80% reduction in parameter overhead, and ablating DPE (using full experts) worsens all metrics on three datasets, which the authors interpret as the parameter decomposition acting as a regularizer.
Hierarchical placement and two-stage training let experts at different depths handle coarse alignment, semantic identification and boundary refinement, yielding leading results on four public benchmarks. DA-SAM3 replaces FFNs only at depths {L/6, L/4, L/2} and uses a two-stage strategy of expert specialization followed by routing calibration, which the authors liken to a radiologist's coarse-to-fine reading process. In Table 1, DA-SAM3 reaches DSC 85.12 and HD 8.425 on Synapse CT, 90.06 and 12.880 on MMWHS, 77.35 and 5.587 on BTCV, and 91.93 and 1.064 on ACDC; ablation shows a 4.37% DSC gain over the SAM3 baseline on Synapse CT.
Perspective
The work targets researchers and engineering teams who need to adapt concept-conditioned segmentation models to medical imaging under resource constraints, in the setting of cardiac and abdominal CT/MRI multi-organ and multi-structure segmentation, using the official splits of four public datasets for comparability with prior work. Methodologically it offers a reusable path: replace feed-forward blocks at selected depths of the fusion module, drive sparse routing with a joint visual and textual signal, and carry expert parameters as a shared frozen base plus low-rank deltas. For readers, this means follow-up work can explore the design space of replacement locations, expert counts and router inputs without starting from full fine-tuning.
The paper reports over 80% reduction in parameter overhead and roughly a 5% accuracy gain, but the main text does not give absolute trainable-parameter counts, inference latency or measured memory usage, so the concrete magnitude of the efficiency benefit still needs confirmation against the code repository. Ablations cover Synapse CT, MMWHS and ACDC, while BTCV does not appear in the ablation table; the choices of placement depths {L/6, L/4, L/2}, four experts and top-k of 2 are mainly empirical, and their sensitivity remains to be observed. In addition, experiments concentrate on cardiac and abdominal CT/MRI, so behavior on other modalities and pathologies is an open question; the text mentions interpretable results but does not elaborate their form in the main body, so readers interested in interpretability would need to consult Figure 2 and the code.
