Skip to main content
Back to timeline
arXivSource publication:

Affine post-processing cuts Wasserstein distance by up to two orders of magnitude for diffusion samplers on high-dimensional and multimodal targets

Related research and updates

Synopsis

The work proposes affine post-processing: splitting a fixed score-estimation compute budget across multiple signal levels, querying the base estimator at each, and pooling the estimates by a linear combination (ridge regression / linear smoother); it proves that at matched compute this split-and-pool estimate has lower score error than a single on-policy estimate, and reports up to two orders of magnitude smaller terminal Wasserstein distance on ill-conditioned Gaussians, multimodal mixtures, and high-dimensional targets.

Source-provided article image: Improving score-based sampling via affine post-processing
Figure 1 ·

Figure 1: (a),(b) Mean W2 for an Ill-conditioned Gaussian against the condition number κ \kappa . Post-processing improves dependency against dimension and κ \kappa . (c), (d) The base estimators generate scattered samples projected onto two dimensions, our approach recovers the true distribution.

arXiv

Interpretation

The authors express the score estimation error as a convex quadratic in the smoothing weights and give a closed-form optimum, called the Wiener oracle post-processor. Prior data-free score-based diffusion spent the whole budget on a single on-policy query; here estimates from multiple signal levels are pooled together with the known terminal values, and the optimal weights are given explicitly. Theorem 1 gives the convex quadratic and its closed-form minimizer; the minimum is unique when the matrix is invertible, otherwise the Moore–Penrose pseudo-inverse gives the minimum-norm solution.

The authors show pure on-policy estimation is optimal among affine post-processors only under a specific zero-gradient condition, which any random unbiased estimator with non-zero variance violates. This theoretically establishes that spending the entire fixed budget on a single query point is almost never the best use of that budget for score estimation. Corollary 1 gives the gradient conditions of the error functional at the on-policy estimator; Corollary 2 proves that under unbiasedness, any non-zero variance makes on-policy estimation non-optimal.

The authors construct a practical non-oracle post-processor via ridge regression in the Legendre polynomial basis, and in experiments its score error captures most of the improvement available to the Wiener oracle. The Wiener oracle requires knowledge of the true score and population bias-variance quantities and cannot be computed; the ridge version relies only on observable query values and variance estimates from repeated queries. Theorem 2 gives uniqueness and equivalence between the constrained ridge regression and its adjoint problem; Figures 2 and 4 report post-processed score error close to the Wiener oracle curves.

In compute-matched experiments, post-processing lowers score error and improves sample quality on ill-conditioned Gaussians, geometry targets, and multimodal mixtures, with Wasserstein distance reduced by up to two orders of magnitude at high dimensions. Base estimators (IS, RDMC) degrade quickly with dimension, while post-processing lets the same estimator match classical chains at equal gradient-evaluation budget. Experiments use IS and RDMC as base estimators and compare against LMC, NUTS, and PT at matched gradient evaluations, reporting seed-averaged Wasserstein distance with min/max error bars.

Perspective

The method applies to data-free score-based diffusion sampling settings where the target log-density is accessible, such as sampling tasks in Bayesian inference, statistical physics, and computational biology. It is designed as a step layered on any base score estimator, and experiments cover ill-conditioned Gaussians, a banana-shaped twisted Gaussian, an orthogonal-covariance mixture, an ellipsoidal shell, eight-mode and forty-mode mixtures, and a separated anisotropic mixture, compared against LMC, NUTS, and PT at matched gradient-evaluation budgets. For researchers and engineers who want to use diffusion samplers on high-dimensional or ill-conditioned targets, this means lower score error and better sample quality can be obtained by splitting the budget and pooling queries across signal levels without changing the base estimator.

When the base estimator explores the space poorly, for example RDMC showing metastability on the eight-mode mixture, post-processing cannot improve generation because off-policy queries supply little useful information; the optimal rule then collapses to pure on-policy estimation, and spending budget on off-policy queries is not worthwhile. As mode separation increases, the Wiener oracle estimate also collapses to the on-policy estimator, and full-budget on-policy estimation is slightly better. The practical implementation relies on estimating the covariance of the score estimator; ablations show cross-level correlations add nothing, and a crude variance estimate suffices to set the weights. Future directions include adaptively choosing design levels and deriving theoretical bounds on score estimation error for the practical non-oracle rule.

Sources