Treating missing modalities as uncertainty: SIUM reaches average Dice of 88.07/87.10/62.52 on BraTS 2018/2020 under missing-modality settings
Synopsis
The work proposes SIUM, which models each modality subset's task representation as a Gaussian (mean carrying task information, variance measuring uncertainty from missing evidence), aligns subset means toward a full-modality anchor via an uncertainty-aware alignment loss that scales variance with their discrepancy, and adds an uncertainty ordering loss keeping subset variance above its supersets, achieving better segmentation than RFNet, mmFormer, M3AE, and DC-Seg across diverse missing-modality configurations on BraTS 2018 and 2020.
Interpretation
A probabilistic representation framework: the task embedding becomes a Gaussian embedding whose mean encodes task information and whose variance encodes intrinsic uncertainty from missing modalities; training samples r=μ+ε⊙σ via the reparameterization trick, and inference uses only the mean. Prior missing-modality methods such as RFNet, mmFormer, M3AE, and DC-Seg encode each modality configuration into a deterministic embedding without explicitly expressing informational incompleteness; this work makes uncertainty an explicit learnable quantity. The paper gives the reparameterization sampling equation and the inference-time mean-only setting, and reports an ablation against the deterministic RFNet baseline on BraTS 2020 (WT 86.98→87.54, TC 78.23→79.20, ET 61.47→63.74).
A theoretical analysis of how uncertainty affects optimization: the gradient covariance Cov_ε[∇_θL] scales with σ², and optimizing the task loss alone drives σ→0 (variance collapse), motivating additional regularization. It turns 'why variance needs regularization' from an empirical practice into a derivable result that then motivates the two loss terms. The derivation uses a first-order Taylor expansion and a covariance identity; a controlled experiment fixes μ at an arbitrary element of r while varying σ, sampling ε 30 times per σ, and shows sampled gradients become increasingly dispersed relative to the deterministic gradient (ε=0) as σ² grows.
Set-inclusive uncertainty guidance: L_UA uses a Gaussian negative log-likelihood to align subset means to a gradient-stopped full-modality anchor while letting σ grow with their discrepancy, and L_UO uses max(0, σ_sup²−σ_sub²) to enforce that subset uncertainty is not lower than its superset. It encodes the nesting structure of modality subsets (S_sub⊂S_sup⊂S_full) directly as an ordering constraint among uncertainties, rather than only using modality masking augmentation. Ablation shows probabilistic embeddings with set-inclusive masking alone give WT/TC/ET of 87.54/79.20/63.74, adding L_UA raises this to 87.73/79.32/65.57, L_UO alone gives 87.57/79.31/64.61, and combining both reaches 88.07/80.22/66.02.
Better segmentation across many missing-modality configurations on BraTS 2018 and 2020, with learned variance inversely related to performance and supersets showing lower variance than subsets. It turns uncertainty from an internal quantity into a testable behavior: variance magnitude corresponds inversely to test Dice and follows the expected set-inclusion ordering. BraTS 2020 average DSC is WT 88.07, TC 80.22, ET 66.02, above RFNet (86.98/78.23/61.47), mmFormer (86.49/76.06/63.19), M3AE (86.90/79.10/61.70), and DC-Seg (87.54/79.63/65.00); BraTS 2018 averages are 87.10/78.37/62.52, above RFNet 85.49/76.00/58.44, M3AE 85.82/77.37/59.85, DC-Seg* 86.59/77.14/59.55, and mmFormer* 86.20/73.65/56.20.
Perspective
The result targets brain tumor segmentation with four MRI modalities (T1, T1c, T2, FLAIR) as input and WT/TC/ET as predicted subregions, with training and evaluation on BraTS 2018 (285 subjects, 199:29:57 split) and BraTS 2020 (369 subjects, 219:50:100 split); volumes are skull-stripped, aligned to a common template, resampled to 1mm³ isotropic resolution, and normalized to zero mean and unit variance, with 112×112×112 patches randomly extracted during training and RFNet's sliding-window strategy with averaged overlapping predictions at inference. The audience that can directly build on it is multimodal medical imaging researchers and practitioners who need stable segmentation when modalities are missing; the authors release code and model checkpoints (https://github.com/atlas-sky/SIUM), enabling reproduction and comparison under the same backbone and training protocol.
Uncertainty values are currently presented as the mean squared magnitude of σ (||σ²||) and as visualizations, with no validation yet of using them as calibrated probabilities for clinical threshold decisions; the inverse variance-performance relationship comes from configuration-level aggregation on BraTS 2020, and whether the same ordering holds on other cohorts, tumor types, or modality combinations still needs external data. The set-inclusive design depends on an enumerable 'subset⊂superset⊂full' hierarchy, so how to define the ordering constraint when the modality set is larger or missing patterns are not nested remains an open question. In addition, the main quantitative results live in tables and figures, so reading only the prose without the tables and figures makes it hard to check per-configuration differences.
