Skip to main content
Back to timeline
arXivSource publication:

SIMCA calibrates guidance weights offline by least squares, letting diffusion and flow-matching posterior sampling match or beat prior methods within 50 steps

Related research and updates

Synopsis

The authors introduce SIMCA: observing that the conditional denoising score-matching objective for diffusion models and the conditional flow-matching objective for flow-matching models are least-squares objectives in the guidance weights at each time step, they reduce guidance-weight tuning in training-free posterior sampling to a two-dimensional linear least-squares problem solved offline by a greedy simulation-based calibration that needs only a single minibatch of sampling trajectories and no retraining; across several inverse problems on FFHQ, ImageNet, CelebA and AFHQ-Cat, the method matches or surpasses state-of-the-art training-free methods and lets diffusion samplers cut steps from 1000 to 50 without significant degradation in reconstruction quality.

Interpretation

The paper observes that when the conditional prediction is written as a weighted sum of the unconditional network output and a measurement-guidance term, the conditional denoising score-matching objective for diffusion and the conditional flow-matching objective for flow matching are linear least-squares objectives in the guidance weights, so only two scalars are fitted per time step. Prior training-free posterior samplers set these weights heuristically or derive them from an assumed posterior covariance; the paper states that none fits the weight to a criterion measurable from data, and all fix the weight of the pretrained field to one. The claim rests on an algebraic derivation of the objective (Eqs. 10 and 11). The paper also notes that because only two parameters are fitted and neither the measurement-consistency term nor the unconditional network is necessarily optimal, the simulation-free calibration does not guarantee sampling trajectories closer to the true posterior.

The paper proposes SIMCA, a greedy simulation-based calibration that replaces the forward-process density in the simulation-free objective with the weight-dependent approximate posterior and calibrates weights sequentially in time, keeping earlier weights fixed, so each step reduces to a linear system. Unlike the simulation-free variant, SIMCA accounts for the states actually visited during sampling, addressing the mismatch between forward marginals and inference-time states that the appendix links to exposure bias in diffusion models. On FFHQ random inpainting, DPS goes from LPIPS 0.203 and FID 69.20 uncalibrated to 0.103 and 33.71 with simulation-free calibration and 0.069 and 19.16 with SIMCA; the same ordering holds for GDM and TMPD. A calibration-set ablation shows a single image suffices, with one weight pair solved in 0.11 s from one image and 6.76 s from one hundred.

Across diffusion and flow-matching backbones, SIMCA matches or surpasses state-of-the-art training-free methods on multiple datasets and inverse problems. The paper reports gains over the strongest baseline DAPS on FFHQ and ImageNet random inpainting, super-resolution and motion deblurring, and on CelebA attains the best value in fourteen of fifteen columns, improving on the strongest baseline Flower on every task. The diffusion evaluation uses 100 held-out images per dataset on FFHQ and ImageNet, with baseline metrics taken from the respective benchmarks under the same operators, noise levels and test images, with hyperparameters grid-searched to maximize PSNR; the flow-matching evaluation follows the five tasks and operators of the Flower benchmark. The paper also reports where SIMCA does not lead: ImageNet Gaussian deblurring trails DDNM, and on AFHQ-Cat deblurring and denoising PnP-GS, Flower and OT-ODE perform best.

Calibrated weights let diffusion samplers reduce the number of steps from 1000 to 50 without significant degradation in reconstruction quality, and each discretization must be calibrated separately. The paper attributes this low-step usability to recalibrating weights at each discretization, and notes that the weights depend on the measurement operator, noise level, measurement-consistency term, pretrained network, sampler and number of timesteps, so any change requires recalibration. On FFHQ with DDPM at 50 steps, random inpainting reaches PSNR 33.76 and SSIM 0.930 versus 32.91 and 0.916 at 1000 steps, while LPIPS and FID degrade moderately; Appendix E shows DDPM remains usable down to about 20 steps, Euler-Maruyama collapses below 50 steps, and DDIM degrades gracefully.

Perspective

The strategy targets training-free posterior sampling with a linear measurement operator and an additive measurement-consistency term, and suits practitioners who need guidance weights quickly when the measurement setup changes often. Calibration is offline on images disjoint from the test images: FFHQ and ImageNet training images for diffusion, and the Flower calibration split for flow matching. Once calibrated, the weights are fixed for a given measurement operator, noise level, measurement-consistency term, pretrained network, sampler and number of timesteps, and add no inference-time overhead; the paper's practical recommendation is DDPM with recalibrated weights at 50 steps, which gives the best distortion metrics of its table on random inpainting and super-resolution.

The paper reports that SIMCA does not surpass the strongest baselines on ImageNet Gaussian deblurring or on AFHQ-Cat deblurring and denoising, which the authors attribute to a Gaussian likelihood that is genuinely limiting on those operators; how much calibration can help there remains open. Because the weights depend on the sampler and the step count, transferring them across configurations rather than recalibrating forfeits the gains, so each configuration needs its own calibration in practice. The conclusions also sit within the scope of image inverse problems, leaving other data modalities and nonlinear measurement operators to be examined.

Sources