Skip to main content
Back to timeline
arXivSource publication:

HierEM treats each site's prostate lesion contour as a noisy view of a latent clean mask, lifting leave-one-site-out Dice to 27.91%–32.67% across three sites

Synopsis

The study proposes HierEM, a hierarchical expectation-maximization framework that treats each site's observed prostate lesion annotation as a noisy observation of a latent clean lesion mask, alternating between inferring a voxel-wise posterior over that latent mask and training a CNN with the posterior as a soft target while estimating site- and case-level sensitivity and specificity under a logistic-normal hierarchical prior; on three-site data it reaches per-site mean DSC of 29.50%–39.69% in pooled held-out evaluation and 27.91%–32.67% in leave-one-site-out generalization, with statistically significant improvements over comparison methods (p < 0.039) and interpretable per-site label-quality estimates (sensitivity α of 31.5%–47.3% at specificity β ≈ 0.99).

Source-provided article image: Deep EM with Hierarchical Latent Label Modelling for Multi-site Prostate Lesion Segmentation
Fig. 1

Fig. 1. Overview of proposed deep mixed EM framework with hierarchical latent label- quality parameters. The E-step infers a latent clean mask, and the M-step updates the segmentation network and site-specific sensitivity and specificity.

· Page 3

Interpretation

The paper frames multi-site annotation variability as a latent-label problem: each case has only one annotation from its acquisition site, and the observed label Y is modeled as a noisy observation of a latent clean mask G, characterized by site- and case-level sensitivity α and specificity β. The authors explicitly distinguish this from STAPLE: STAPLE addresses multi-annotator label fusion for a fixed image, whereas here there is typically one annotation per case and the annotation source is the acquisition site rather than multiple independent raters, so no label fusion is performed and a site-agnostic lesion representation is learned from single, site-dependent noisy labels. The distinction is stated as a problem-setting argument in the introduction and formalized by the conditional independence assumption in Eq. (1); no direct numerical comparison against STAPLE is reported.

HierEM uses a logistic-normal hierarchical prior to decompose label quality into global, site, and case components, with logit(α)=µα+as+uk and logit(β)=µβ+bs+vk, placing zero-mean Gaussian priors on as, bs, uk, vk (equivalent to L2 penalties under MAP) and constraining site effects to sum to zero for identifiability. Compared with the non-hierarchical Site-EM baseline that estimates independent αs and βs per site, the hierarchical structure enables partial pooling across sites, regularizes site-level deviations, and discourages overfitting to center-specific contouring styles. The decomposition is formalized in Eqs. (2)–(3) and compared against Site-EM in the ablation; the paper reports that the L2 penalty provides shrinkage that stabilizes site/case quality estimates and prevents degenerate sensitivity and specificity values.

Learning alternates an E-step and M-steps: the E-step computes the voxel-wise latent-mask posterior q via Eq. (5), M-step (A) updates the UNet with q as soft labels using a cross-entropy plus Dice loss, and M-step (B) maximizes the L2-penalized Q(ϕ) over aggregated expected sufficient statistics (TPk, Pk, TNk, Nk) with L-BFGS. Because M-step (B) depends on the data only through low-dimensional sufficient statistics, it can be solved efficiently with a few second-order iterations; in practice each EM iteration runs 5 gradient epochs for θ and 5 optimizer steps for ϕ, with auxiliary supervision on observed labels decayed by a sigmoid schedule over the first five epochs to stabilize early optimization. Algorithm 1 lays out the full procedure, and hyperparameters (Adam, learning rate 10⁻⁴, 5 epochs/5 steps, auxiliary-supervision decay) are stated explicitly in the text.

On three-site data, HierEM achieves the best mean Dice on all three sites in the pooled held-out split (Site 1: 39.69%, Site 2: 29.50%, Site 3: 35.60%) and Dice of 28.11%, 27.91%, and 32.67% in the harder leave-one-site-out setting, versus 25.50%, 24.66%, and 31.20% for supervised UNet, with lower HD95; risk-coverage curves show its uncertainty concentrates errors in the rejected region, and it attains the highest sensitivity at matched specificity of about 0.99. The paper reads the gap between pooled and leave-one-site-out performance as evidence of strong domain shift, suggesting that pooled-split performance largely comes from learning site-specific contouring biases, and that explicitly modeling site/case-dependent label noise narrows this gap. Results come from three-site datasets (pooled held-out sets of 252, 87, and 265 cases for Sites 1–3; LOSO sets of 1201, 391, and 1265 cases), analyzed with paired t-tests and Benjamini-Hochberg FDR at 5%; the paper also notes that not every comparison is significant in the pooled split, likely due to limited test-set size.

Perspective

The work targets multi-site prostate lesion segmentation where each case has a single site-dependent annotation and where site-specific fine-tuning or calibration at deployment is not feasible or practical; the method is described as backbone-agnostic, so the segmentation network can in principle be swapped. It also outputs site-level sensitivity and specificity as diagnostics of annotation behavior, which the authors stress represent pooled site-level labeling tendencies rather than individual-reader performance. Stated future directions include extending to multi-site, multi-annotator datasets and richer models of clinical annotation variability.

The paper reports that not all comparisons reach significance in the pooled held-out split, which the authors attribute to limited test-set size, so the margin of advantage in that setting still needs confirmation on larger test sets. Leave-one-site-out results rest on three sites, so how stably the hierarchical prior estimates site effects with more sites remains to be seen. Sensitivity and specificity are interpreted as pooled site-level labeling behavior rather than individual-reader performance, and their relationship to true clinical annotation quality remains an open question. In addition, the risk-coverage curves and sensitivity comparison are presented in Figure 2, whose specific values are not included in the available text, so these uncertainty metrics can only be understood as described in the prose.

Sources