SEA-PEFT raises mean Dice by 2.4–2.8 points in 1/5/10-shot 3D medical segmentation while training under 1% of parameters
Synopsis
The authors propose SEA-PEFT, which treats adapter configuration as an online allocation problem solved during fine-tuning through a search–audit–allocate loop that trains active adapters, estimates each adapter's Dice utility by momentarily toggling it off, and reselects the active set under a parameter budget with a greedy knapsack allocator, stabilized by EMA+IQR smoothing and a finite-state machine; on TotalSegmentator and FLARE'22 it improves mean Dice by 2.4–2.8 points over the strongest fixed-topology PEFT baselines across 1/5/10-shot settings while training under 1% of parameters.
Interpretation
SEA-PEFT moves the choice of adapter type, insertion site, and rank from pre-training manual or offline search to online decisions during fine-tuning: the search phase trains currently active adapters, the audit phase runs on/off perturbations on a sampled subset, and the allocate phase greedily activates adapters under a parameter budget. Prior PEFT methods such as FSEFT and FreqFit use fixed single-adapter designs, while offline search methods such as NOAH, AutoPEFT, and Fairtune require fully fine-tuning each candidate configuration; SEA-PEFT selects the configuration within a single run without an offline sweep. The paper presents the algorithm and objective (a knapsack form in Eq. 4) and reports Dice on TotalSegmentator and FLARE'22 across 1/5/10-shot settings, averaged over 3 random seeds on a hold-out test set.
Directly measuring each adapter's marginal Dice contribution via on/off perturbation, then suppressing estimation jitter in few-shot noise with a two-stage EMA+IQR filter and a finite-state-machine voting mechanism. AdaLoRA reallocates LoRA rank during training based on singular value magnitudes in weight space, whereas SEA-PEFT's utility estimate is directly aligned with the segmentation metric Dice and is model-agnostic; EMA+IQR and the FSM are stabilization components aimed at high-noise few-shot regimes. The paper provides Proposition 1 on EMA bias and variance bounds and IQR concentration, Lemma 1 on an audit coverage lower bound, and Proposition 2 on FSM structural-error reduction; the implementation uses 30% active / 70% inactive sampling plus an ε-exploration term.
On TotalSegmentator, SEA-PEFT attains 78.33/79.90/80.29 Dice in the 1/5/10-shot settings, and on FLARE'22 it attains 75.77 (1-shot) and 76.33 (5-shot), above fixed-topology PEFT baselines. Mean Dice improves by 2.4–2.8 points over the strongest fixed-topology baseline; on the SuPreM-pretrained backbone it reaches 77.89 average at 5-shot and 79.93 at 10-shot, with the margin over the best fixed baseline widening from 0.88 points at 5-shot to 3.58 points at 10-shot. Results come from Tables 1 and 2, covering nine abdominal organs, binary and multi-class tasks, and two pretrained weight sets (FSEFT and SuPreM), all evaluated on frozen Swin-UNETR backbones.
End-to-end search–audit–allocate plus final re-fine-tuning takes roughly 2.5, 4.5, and 6.5 hours in the 1/5/10-shot settings while consistently training only 0.2% of parameters. The paper notes that offline search methods would require hundreds of 3D training runs over the audit space at 1–10 shots, which the authors consider infeasible to compare directly; SEA-PEFT's online design evaluates only a small subset of adapters per audit cycle, so compute scales with data rather than configuration-space size. Times and parameter fractions come from Table 3; experiments run on a single V100 32GB GPU with 96×96×96 patches, batch size 1, AdamW, learning rate 1e-3, and 200 epochs with early stopping.
Perspective
The work targets few-shot 3D medical image segmentation, specifically 1/5/10-shot adaptation on frozen Swin-UNETR backbones with FSEFT and SuPreM pretrained weights, on nine abdominal organs in TotalSegmentator (binary) and FLARE'22 (multi-class). It enables clinical teams without dedicated AI engineers to potentially complete site adaptation within hours, because configuration selection happens online within a single run and only 0.2% of parameters are trained. It assumes an available validation set to supply Dice feedback and an enumerable library of adapter candidates (LoRA, AdaptFormer, Affine-LN with SA/PA/SAPA topologies).
The paper notes that NOAH and AutoPEFT require fully fine-tuning each candidate configuration, which over the audit space at 1–10 shots would require hundreds of 3D training runs, so no direct comparison is made; readers judging the advantage over offline search would need to assess this comparison gap themselves. The stabilization effects of EMA+IQR and the FSM are presented through theoretical propositions and overall Dice, and the sensitivity of hyperparameters such as β, λS, τact, τrank, and µeff is not expanded in the main text. In addition, results concentrate on abdominal organs and two pretrained backbones, so behavior on other anatomical regions, imaging modalities, or backbones remains an open question.
