Decoder-side multi-kernel gated adapters raise CNN TI-RADS diagnostic accuracy from 0.406 to 0.632 and improve external segmentation Dice under cross-center thyroid ultrasound shift
Synopsis
Training a unified multi-task model on ThyroidXL (11,635 images, 4,093 patients) and testing externally on DDTI (660 images), this work characterizes negative transfer between segmentation and TI-RADS malignancy classification under cross-center domain shift for a CNN (ResNet34) and a medical ViT (MedSAM), and proposes lightweight decoder-side adapters, MKGA and its residual variant ResMKGA, which refine multi-scale skip features with complementary receptive fields and apply semantic context-conditioned gating to suppress artifact-prone content; the adapters improve out-of-domain segmentation stability (ResNet34+MKGA external Dice 0.659, ResMKGA 0.671, versus 0.590 for the unfrozen baseline) and, in the CNN setting, significantly raise TI-RADS diagnostic accuracy (0.406 to 0.
Interpretation
The study characterizes negative transfer in multi-task thyroid ultrasound under cross-center shift: geometry-driven segmentation and texture-driven TI-RADS malignancy assessment degrade differently when forced through a shared encoder. Most prior multi-task pipelines implicitly assume a single shared backbone can serve both objectives; this work empirically shows that assumption is brittle under cross-center shift by evaluating both ResNet34 and MedSAM backbones. Compares two backbones on the in-domain ThyroidXL test set and the external DDTI test set, reporting segmentation Dice, TI-RADS Acc/F1/AUC, and positioning metrics, with Wilcoxon, McNemar, and DeLong tests under Benjamini-Hochberg FDR correction.
It proposes lightweight decoder-side adapters, MKGA and its residual variant ResMKGA, which refine multi-scale skip features with complementary receptive fields and gate skip content conditioned on semantic context before fusion. Unlike approaches that rely on a shared encoder or optimization-only fixes such as PCGrad, this method places robustness in the decoder's skip-fusion stage as a parameter-efficient structural solution. Ablations on ResNet34 isolate gating, multi-kernel refinement, kernel size, and SE channel recalibration, reporting p-values (e.g., removing the gate lowers external TI-RADS AUC from 0.642 to 0.589).
The adapters improve cross-center segmentation stability and, in the CNN setting, significantly improve TI-RADS diagnostic decisions. Unfrozen ResNet34 drops from 0.861 in-domain Dice to 0.590 external, while adding MKGA/ResMKGA raises external Dice to 0.659/0.671 (p<0.05); TI-RADS diagnostic accuracy rises from 0.406 to 0.632 (McNemar p<0.001). Based on the external DDTI test set with statistical testing; however, the AUC gain is not significant (DeLong p=0.45), so the evidence centers on diagnostic decision accuracy rather than ranking ability.
The study observes complementary transfer behavior between ViT and CNN: MedSAM transfers well for segmentation, while ResNet34 is more reliable for malignancy classification and positioning. This provides task-dependent empirical guidance for backbone selection, rather than assuming medical foundation models are uniformly better across tasks. MedSAM+ResMKGA+LoRA (r=4) achieves the best external Dice (0.675), but its advantage over lightweight ResNet34 variants is not significant (p>0.05); MedSAM variants show external malignancy AUC around 0.48-0.50.
Perspective
This work targets thyroid ultrasound scenarios that require simultaneous nodule segmentation and TI-RADS risk stratification under cross-center, artifact-laden conditions (calipers, text overlays), and is suited to researchers and engineering teams seeking parameter-efficient improvements in out-of-domain robustness. Its conclusions rest on training with ThyroidXL and external testing on DDTI, with positioning evaluated only on ThyroidXL (DDTI lacks position labels), so the results mainly apply to multi-task ultrasound pipelines with similar annotation structure.
The external test set DDTI has 660 images and is used only for testing, so whether the conclusions generalize to more centers, devices, and larger samples remains to be verified; the TI-RADS AUC gain is not statistically significant (DeLong p=0.45), meaning evidence for improved diagnostic ranking is weaker than for accuracy; the ViT's external malignancy performance (AUC around 0.48-0.50) suggests texture cues remain unstable under strong shift, and achieving consistent cross-center diagnostic robustness across backbones remains an open question.
