LightMIS reaches 86.71% modality-macro Dice across six medical segmentation datasets with 0.131M parameters and full GPU delegation on a smartphone
Synopsis
The work introduces LightMIS, a family of ultra-lightweight convolutional networks that aligns the outputs of a five-level encoder to a common resolution with Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade, thereby removing the learned stage-wise decoder; evaluated by five-fold cross-validation under a common nnU-Net v2.3.1 protocol on DRIVE, Kvasir-SEG, DSB18, BUSI, ISIC-2017, and ISIC-2018, the full model reaches 86.71% modality-macro Dice and 78.99% IoU with 0.131M parameters and 0.575 GFLOPs, within 0.04 and 0.08 percentage points of Mobile U-ViT's 86.75% and 79.07% while using 90.58-99.61% fewer parameters and 82.54-96.
Interpretation
It proposes Progressive Receptive Fusion (PRF), a lightweight module that enriches narrow feature representations through temporary channel expansion, complementary depthwise receptive fields, and progressive cross-branch information transfer. Relative to the Cascade Multi-Receptive Fields block in TinyU-Net, PRF differs by propagating a channel-averaged spatial response from each processed branch to the subsequent branch before convolution, so information transfer does not require matching channel dimensions between branches. PRF ablations on BUSI and Kvasir-SEG show that receptive-field diversity improves all reported metrics over uniform depthwise processing, that progressive branch interaction gives a substantial BUSI overlap improvement while its Kvasir-SEG overlap effect is small and boundary metrics slightly favor the non-interacting variant, and that in parameter-matched encoder replacements PRF attains the highest macro Dice with leading scores on BUSI and DRIVE.
It introduces the Adaptive Fusion Cascade (AFC), which sequentially combines the Adaptive Kernel Fusion (AKF) block adopted from AULUNet with its final ReLU replaced by GELU, the proposed PRF module, and residual adaptive-kernel refinement, applied at every encoder level and in the final aggregation head. Relative to single-AKF or single-PRF configurations, the complete AKF-PRF-ResAKF sequence obtains the highest BUSI Dice and IoU, and replacing the final AKF with ResAKF raises modality-macro Dice from 85.74% to 85.92%. The authors note that these composition comparisons differ in parameter count and therefore do not isolate composition independently of model capacity, and they add parameter-matched encoder and prediction-head comparisons; the ResAKF gain is not uniform, with small reductions on the two ISIC datasets and DSB18, so it is retained for its aggregate trade-off.
It develops LightMIS, a direct-aggregation architecture in which Scale-Aligned Projection (SAP) blocks project all five encoder outputs to a shared channel width and bilinearly align them to a common spatial resolution, after which the concatenated features are projected and refined by a final AFC and passed through a lightweight prediction head whose logits are bilinearly resized to the input resolution, with no symmetric multi-stage decoding path. Unlike decoder-free works such as FSE-Net, CBL-Net, and FAN-Net, whose evaluation remains largely focused on retinal vessel segmentation, LightMIS is evaluated without task-specific modifications under a unified nnU-Net protocol across six datasets and five imaging settings. In parameter-matched head comparisons, SAP with a final AFC reaches 79.62% Dice on BUSI, 87.52% on Kvasir-SEG, and 81.76% on DRIVE for a macro of 82.97%, above FPN-lite at 82.56% and a parameter-matched U-Net decoder at 82.38%; a deepest-feature-only configuration drops to 52.62% Dice on DRIVE, indicating the importance of early encoder features for thin retinal structures.
It provides a unified evaluation under the nnU-Net v2.3.1 protocol with five-fold cross-validation across six datasets, together with desktop and smartphone complexity, latency, and memory measurements. LightMIS obtains the second-highest modality-macro Dice and IoU and the best mean Dice rank across the six datasets, while reducing parameter count by 90.58-99.61% and GFLOPs by 82.54-96.14% relative to Mobile U-ViT, nnWNet, and nnU-Net. Paired bootstrap and permutation tests show that among comparisons with nnU-Net, nnWNet, and Mobile U-ViT, LightMIS achieves statistically confirmed superiority only on DRIVE, where the confidence interval lies entirely above zero and the adjusted p-value is significant; the authors explicitly describe the macro-average differences as numerically close rather than as evidence of statistical superiority. On the smartphone all variants achieve 100% GPU delegation, but the authors also note that GFLOPs alone do not determine real-device efficiency: UNeXt and full LightMIS have nearly identical computational costs of 0.577 and 0.575 GFLOPs yet run at 42.14 ms versus 138.43 ms on the Mali-G52.
Perspective
The results target public 2D binary medical image segmentation, covering five imaging modalities: retinal vessels, gastrointestinal polyps, cell nuclei, breast lesions, and skin lesions, and suit settings where a segmentation model must run on mobile or resource-constrained devices. The authors release the implementation, environment specification, exact five-fold partitions, nnU-Net plan files, dataset-conversion scripts, baseline commit hashes, training commands, checkpoints, metric implementation, FLOP-counting script, LiteRT conversion code, and desktop and mobile benchmarking scripts, enabling reproduction and extension under the same protocol. Future directions include patient-grouped and external validation, multiclass and volumetric segmentation, quantized execution on additional edge processors, and generalizing PRF and AFC to other dense prediction tasks.
For DRIVE, ISIC-2017, and ISIC-2018, publicly labelled official partitions were pooled before cross-validation, so the reported numbers are pooled-label cross-validation results rather than official challenge-test results; cross-validation was performed at image level without subject grouping or category-based stratification, so independence across folds cannot be guaranteed where multiple images may originate from the same subject, lesion, or acquisition sequence. Except for DRIVE, images were directly resized without preserving aspect ratio, which may alter anatomical or lesion geometry. The modality-macro average first combines ISIC-2017 and ISIC-2018 into a single dermoscopy score and then weights the five modalities equally, differences between leading models are small, and the current paired analyses use fixed out-of-fold predictions and do not capture variation from repeated model training. The mobile evaluation covers one smartphone GPU and one software stack, measures delegated model execution rather than complete application latency, and does not address other edge processors, sustained thermal conditions, energy budgets, or clinical workflows. In addition, the unified nnU-Net protocol may not be optimal for baselines designed around pretraining, different optimizers, deep supervision, or model-specific training schedules, and the authors state that it provides a controlled comparison but is not necessarily fair to all architectures.
