CalcSeg combines confidence-aware curriculum learning with slice-wise self-attention to raise myocardial scar segmentation Dice to 0.677 and low-confidence Dice to 0.644 on single-stack LGE-CMR
Synopsis
The work presents CalcSeg, a confidence-aware latent 3D context curriculum learning framework that scores each sample using Dice, percentage scar-burden error, and epistemic uncertainty from Monte Carlo Dropout, expands training from easy to hard cases across stages, and uses slice-wise self-attention to infer subject-level 3D anatomical context from single-stack 2D LGE-CMR; on LGE-CMR data from four sites and two segmentation challenges (MICCAI 2012 LV Infarct, EMIDEC 2020) comprising 976 patients, it reaches myocardial scar Dice of 0.677±0.24, low-confidence Dice of 0.644±0.22, scar error of 38.88%, and low-confidence scar error of 35.03%, outperforming TransUNet, AttentionUNet, UNETR, ScarNet, and ScarNet with supervised curriculum learning using expert difficulty labels.
ule (as shown on the left panel of Fig. 1) that explicitly models subject-level 3D representations inferred from slice-wise 2D features, z; thereby improving seg- mentation across sparse LGE stacks. In contrast to previous approaches that pro- cess individual 2D slices [31], our model learns hierarchical inter-slice context in the latent space. More specifically, we develop an autoencoder architecture that integrates a Transformer-based encoder [8, 23] with a multi-scale convolutional decoder enhanced by a disentangled slice-wise self-attention [1]. This module is designed to effectively capture subject-level anatomical context by performing explicit 2D-to-3D latent feature fusion on the encoded latent features, z. In order to preserve slice order information, fixed sinusoidal positional embeddings are integrated along the slice dimension prior to attention modeling.
· Page 4Interpretation
It introduces a semi-supervised, model-adaptive sample-difficulty definition that combines predictive inaccuracy (Dice ρ and percentage scar-burden error v) with predictive uncertainty (epistemic uncertainty ϵ estimated by Monte Carlo Dropout) into a difficulty vector Di=[γ1ρi, γ2vi, γ3ϵi], and decides whether a sample enters the current training subset via stage-dependent thresholds σk computed automatically from the median of scores at each stage. Existing curriculum learning typically relies on static predefined difficulty definitions, sample reweighting, or local-to-global schedules and does not explicitly model subject-level variability or pathology-specific progression; here difficulty is grounded in the network's own training dynamics and predictive behavior without manual difficulty labels. The confidence scoring function and stage-threshold rule are given as equations in the method section and are compared in ablation against slice attention alone and against ScarNet with supervised curriculum learning using expert difficulty labels.
It proposes latent 3D context modeling: a Transformer encoder (weights loaded from MedSAM) with a multi-scale convolutional decoder, applying single-head slice-wise self-attention on latent features z as A=softmax(QKᵀ/√C)V, with fixed sinusoidal positional embeddings and a binary mask that suppresses padded slices to support variable stack sizes. Prior approaches process isolated 2D slices or introduce only partial inter-slice information without modeling a complete 3D anatomical relationship; this module performs explicit 2D-to-3D latent feature fusion while preserving slice-specific features. Ablation shows that adding slice-wise self-attention alone raises Dice from 0.571±0.24 to 0.639±0.23 and lowers scar error from 53.68% to 47.46%.
It implements sample-space continuation via homotopy-based optimization, with a total loss that is a weighted sum of focal and foreground-weighted Dice loss, where wfg varies from 0.6 to 0.8, β from 0.25 to 0.35, and δ from 1.25 to 2 across stages, moving training from stable coarse learning to fine-grained refinement. It interprets curriculum learning as continuation optimization over the full objective, reshaping the loss landscape by progressively introducing lower-confidence samples rather than relying on random sampling or fixed difficulty. The loss form and parameter intervals are stated explicitly in the method section, with parameters set on fixed intervals based on the difficulty score Di for ease of optimization.
It validates on multi-center data: 976 patients with ischemic and non-ischemic cardiomyopathies and without fibrosis, 4 to 18 slices per subject, with a 7:1:2 train/validation/test split, where CalcSeg outperforms five baselines overall and on clinically challenging cases, and Monte Carlo uncertainty for the clinically challenging subgroup decreases substantially across curriculum stages. Relative to the current SOTA ScarNet with supervised curriculum learning, low-confidence Dice rises from 0.574±0.23 to 0.644±0.22 and low-confidence scar error falls from 48.34% to 35.03%. Results come from the complete independent test set and 6 known challenging cases explicitly flagged by clinical experts, reported with means and standard deviations; the overall mean uncertainty appears relatively stable because high-confidence cases dominate the test set (ratio 22:6).
Perspective
The result targets the single-stack LGE-CMR acquisition setting, for patients with ischemic and non-ischemic cardiomyopathies and without fibrosis, with 4 to 18 slices per subject, where slices are masked to the LV region and cropped and resized to 224²; the method initializes the Transformer encoder with MedSAM weights and pretrains the autoencoder on stacks of 2D LGE slices, runs on NVIDIA A40 and V100 GPUs, and uses 100 Monte Carlo Dropout samples per subject to quantify uncertainty. The authors state that future work will include local-to-global curriculum approaches, addressing uncertainty in scar tissue boundary, and incorporating multi-sequence CMR for better inflammation detection.
The overall mean uncertainty appears relatively stable because high-confidence cases dominate the test set (ratio 22:6), while uncertainty decreases more clearly for the challenging subgroup, so stratified evaluation is key to interpreting the results; difficulty thresholds are determined automatically from the median of scores at each stage, and the weights γj and stage parameters are set on fixed intervals, whose influence on results is not explored in the text; in addition, uncertainty in scar tissue boundary is listed by the authors as future work, indicating that boundary performance remains an open question.
