L2L-Flow moves volumetric stochastic segmentation into latent space: about 14x faster radiotherapy-target inference with a competitive GED of 0.161
Synopsis
The work introduces Latent-to-Latent Flow (L2L-Flow), which first compresses labels into a latent space with a volumetric label autoencoder and then learns a rectified flow between an image-conditional latent prior and frozen latent label representations, yielding stochastic segmentation on a private radiotherapy clinical-target-volume dataset (55 cases, 5-fold cross-validation) and the CURVAS multi-organ dataset (90 cases), with inference about 14x faster than full-resolution Flow-SSN (0.936 s/image versus 15.296 s/image) at a GED of 0.161, alongside a time-shifted noise schedule that improves Flow-SSN stability on high-dimensional volumetric data.
Interpretation
The authors identify performance degradation of Flow-SSN on high-resolution volumetric medical data and propose a time-shifted noise schedule yt=(1−ts)u+tsy with ts=t/(α−t(α−1)), where larger α induces more aggressive noising. The noise schedule of Flow-SSN had not been calibrated for high-dimensional volumetric tasks; the authors adapt the shift schedule from modern flow models to a path direction running from t=0 to t=1 and provide an ablation over α. On the radiotherapy dataset under 5-fold cross-validation, the shifted Flow-SSN lowers GED from 0.165±.008 to 0.155±.005 and raises Dice from 0.894±.004 to 0.899±.001; in the ablation α=3 performs best, while larger α is more prone to overfitting.
The authors propose L2L-Flow: train a volumetric label autoencoder to obtain compressed label latents z1=Eϕ(y), then learn a linear flow interpolant from an image-conditional latent prior N(Eλ(x),I) to frozen latent label representations, trained with a least-squares regression objective. Unlike Flow-SSN, which models the joint in observation space, this method places both source and target in an encoded latent space, making volumetric sampling computationally feasible; unlike standard latent diffusion that models a single image variable, it models the label conditional distribution. On the radiotherapy dataset L2L-Flow reaches GED 0.161±.004, Dice 0.896±.002, diversity 0.182±.008, 9.3M parameters, and 0.936±.016 s/image, compared with 15.296±.013 s/image for Flow-SSN.
L2L-Flow offers a controllable efficiency/performance trade-off: 4x downsampling to a 48³ latent resolution on radiotherapy, and an additional 2x downsampling to 96³ on CURVAS. The authors expose the downsampling factor as an explicit control knob and report that the 2x variant narrows the gap to full-resolution models on the pancreas at the cost of slower inference. On CURVAS the 2x model remains about 5x faster than full-resolution Flow-SSN; the 4x model shows a larger performance gap on the already small pancreas due to compression, with Fig. 5 marking the part of the pancreas it fails to segment precisely.
Qualitative comparison shows that samples from the flow-based models (L2L-Flow and Flow-SSN) vary in a way that better resembles the IOV contours of 10 clinical experts, whereas SSN contours vary only in a strong concentric manner. The authors note this clinical-relevance distinction is not captured by the GED metric alone, arguing for combining quantitative evaluation with qualitative and clinical considerations in clinical workflows. The conclusion comes from a qualitative evaluation of one case with ground-truth IOV from 10 experts, and the authors state explicitly that GED alone does not capture this distinction.
Perspective
The work targets stochastic segmentation uncertainty arising from inter-observer variability in volumetric medical images, suited to downstream planning settings such as radiotherapy clinical target volume delineation and multi-organ segmentation where contour variability must be simulated; the efficiency/performance knobs the authors provide (4x downsampling to 48³ on radiotherapy, 2x downsampling to 96³ on CURVAS) let users pick a configuration by compute budget and structure size. In the conclusion the authors propose future work integrating more expressive image and label encoders, such as those derived from foundation models, to further enhance L2L-Flow.
The radiotherapy dataset is private and has only 55 cases, and CURVAS has 90 cases, both evaluated with 5-fold cross-validation, so sample size limits how far the conclusions generalize; the authors also note that moving complex probabilistic models from 2D to 3D involves methodological and computational difficulties, such as SSN's low-rank sampling instability and MoSE's reduced training sample size due to memory constraints. In addition, the authors stress that GED alone does not reflect clinical relevance, so readers interpreting metric rankings should still weigh qualitative samples and clinical judgment; the stronger encoders proposed for future work are not yet validated in this work.
