Skip to main content
Back to timeline
arXivSource publication:

KANResDiff combines KAN spline time encoding with a local Schrödinger bridge for residual diffusion, cutting GED by up to 18.8% and raising HM-IoU32 by up to 7.7% on LIDC and ISIC3

Synopsis

The work proposes KANResDiff, which learns local residual diffusion via Kolmogorov-Arnold Networks for ambiguous medical image segmentation: it replaces MLP linear time embeddings with B-spline-based Independent Time Encoding to strengthen independence across inference stages, and injects a deterministic residual prior with learnable weights through a Residual Schrödinger Bridge, achieving state-of-the-art GED and HM-IoU on the two public datasets LIDC and ISIC3 with maximum improvements of 16.8% and 7.7% respectively while keeping competitive MDM performance.

Source-provided article image: KANResDiff: Learning Local Residual Diffusion via Kolmogorov-Arnold Network for Ambiguous Medical Image Segmentation
Fig. 1

Fig. 1. (a) Traditional MLP-based time embeddings can only be optimized globally. (b) Our proposed KANResDiff enhances indepedence across timesteps via spline-based time embeddings, which can be locally optimized. (c) Traditional residual diffusion leads to large gap between forward and backward diffusion path, hindering effective semantics modeling. (d) Our proposed KANResDiff constructs local Schrödinger Bridge with dynamic residual components to form a least-costed diffusion path, where forward and backward diffusion paths are close to each other.

· Page 2

Interpretation

It proposes KANResDiff, which turns stochasticity injection for ambiguous medical image segmentation from a fixed, predefined scheme into a dynamic and progressive process that assigns distinct roles to different inference stages. The authors describe it as the first paradigm enabling dynamic and progressive stochastic modeling for ambiguous medical image segmentation, in contrast to cVAE-based, logit-distribution-based, and existing diffusion methods that inject randomness only at the final prediction stage or uniformly along the whole trajectory under fixed stochastic formulations. This is a positioning claim at the method level, supported by the challenge-and-solution contrast in Fig. 1 and the overall framework in Fig. 2, without a quantitative comparison of mechanisms against other paradigms.

It proposes Independent Time Encoding (ITE), which reparameterizes time embeddings with locally supported B-spline basis functions so that optimization at one timestep yields zero gradient for timesteps outside the support interval, enhancing independence across timesteps. Whereas MLP time embeddings share the parameter set W across all timesteps so that optimizing one timestep modifies every embedding, ITE achieves locally independent optimization through the local support property in Eqs. (3)-(5); the implementation uses a cubic B-spline (p = 3) and relies on the KAN's resistance to forgetting and continuous output. The paper provides gradient derivations (Eqs. 1, 2, 4, 5) as the mechanistic argument; the ablation on LIDC (Table 3) shows that adding ITE alone lowers GED100 from 0.187±0.002 to 0.181±0.003, raises HM-IoU32 from 0.667±0.002 to 0.683±0.001, and raises MDM32 from 0.903±0.003 to 0.908±0.002.

It proposes the Residual Schrödinger Bridge (RSB), which constructs a local Schrödinger bridge at each inference stage and uses a dynamic weight hRSB(t) to regulate the contribution of the deterministic residual term, forming a flexible determinism-stochasticity coordination. Traditional residual diffusion introduces residuals through a fixed formulation (Eq. 6) without maintaining the true anatomy of diffusion paths in latent space; RSB instead uses the learnable-weight form in Eq. (7) and treats the residual term as a control drift inside the Schrödinger bridge forward and backward PDEs (Eqs. 8, 9) for minimum-cost path optimization. The paper gives two training objectives for the forward and backward procedures (Eqs. 10, 11), refining a pretrained DDPM with a fixed noise schedule; the ablation (Table 3) shows that adding RSB alone lowers GED100 to 0.174±0.004, raises HM-IoU32 to 0.688±0.004, and raises MDM32 to 0.916±0.002.

On the two public datasets LIDC and ISIC3, KANResDiff achieves the best GED and HM-IoU while keeping MDM comparable to supervised sample-level segmentation methods. The paper reports maximum reductions of 18.8%, 16.6%, and 16.8% on GED16, GED32, and GED100, and a maximum 7.7% improvement in HM-IoU32; on LIDC the values are 0.172±0.001 for GED100, 0.701±0.001 for HM-IoU32, and 0.930±0.002 for MDM32, and on ISIC3 they are 0.139±0.003, 0.773±0.002, and 0.942±0.003. Results are reported as mean ± standard deviation over five runs and compared against Prob. Unet, MoSE, P2SAM, CIMD, AB, CCDM, SSB, ContourMS, SSN, c-Prob. Unet, and c-SSN; the ablation shows further gains when ITE and RSB are used jointly, which the authors read as evidence that the two components are complementary.

Perspective

The work targets ambiguous medical image segmentation settings that require multiple plausible segmentation hypotheses, and applies to public data with multiple annotator labels, such as LIDC lung CT with four segmentation labels and the ISIC3 dermoscopic image subset with three labels. The implementation uses a single RTX 4090, T = 1000, batch size 8, the Adam optimizer with a learning rate of 10^-4, and 600 training epochs for both ITE and RSB, refining a pretrained DDPM with a fixed noise schedule. For researchers and practitioners who want to move away from hand-set stochastic intensity and let early stages emphasize structural consistency while later stages express diversity, the ITE and RSB combination offers a reusable design pattern; the source code is released at the repository given in the paper.

The evaluation centers on the two datasets LIDC and ISIC3 and on the three metric families GED, HM-IoU, and MDM, so behavior under other imaging modalities, other annotation styles, and other evaluation dimensions remains an open question. RSB relies on a pretrained DDPM with a fixed noise schedule, and its behavior under different backbones or noise schedules is worth watching. In addition, the complementarity of ITE and RSB is currently supported by the four-variant ablation on LIDC, and their relative contributions under different data distributions still await further verification.

Sources