PreviewDiff puts a multimodal critic inside the denoising loop, beating Best-of-N at matched compute on SDXL and LTX-Video
Synopsis
PreviewDiff introduces a training-free test-time search method that decodes partial previews at selected denoising checkpoints, asks a multimodal judge to score and critique them, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations that are scored and selectively rolled forward, consistently outperforming budget-matched Best-of-N and scalar-search baselines on image and video generation benchmarks.
Interpretation
It recasts diffusion sampling from scalar search into multimodal critic-guided tree search over intermediate latents: node states are intermediate clean-latent estimates, prompt conditionings, and feedback history, while actions are critic-proposed semantic corrections plus local re-noising and continuation. Best-of-N only selects among completed outputs, and EvoSearch, particle sampling, and Video-T1 use the verifier as a scalar reward or search over noise, particles, or frame candidates; PreviewDiff lets the same multimodal model act both as a value estimator and as a policy-like source of semantic actions, with branches over interpretable object-, attribute-, relation-, and action-level fixes. The paper provides a full derivation, including clean-latent estimates and local re-noising rules for both DDIM-style and flow-matching samplers, plus the selection, expansion, rollout, and backup steps of beam-pruned tree search.
On image and video generation, PreviewDiff improves as the denoiser-step budget grows and stays above budget-matched Best-of-N at every tested budget. On SDXL the mean Gemini score rises from 2.17 to 2.86 while Best-of-N rises from 1.91 to 2.71; on LTX-Video PreviewDiff rises from 1.99 to 2.39 while Best-of-N rises from 1.75 to 2.13, indicating intermediate feedback is not merely a low-budget shortcut. The main scaling experiments use SDXL and LTX-Video, report compute in denoiser steps and verifier calls, and compare against both verifier-call-matched and forward-pass-matched Best-of-N settings.
Fixed-budget transfer across backbones and benchmarks: on images PreviewDiff improves over Best-of-N on most reported metrics across SDXL, SD-3.5, and FLUX.2, and on video it improves over Best-of-N on LTX-Video and Wan 2.2, with gains also appearing under auxiliary evaluators such as Qwen-3, BLIP-VQA, Qwen-2.5, and LLaVA. The paper uses Qwen-3 both as a critic and as an evaluator and reports non-Gemini evaluation columns to argue the gains are not merely overfitting to the Gemini critic. Tables cover GenEval2, T2I-CoReBench, T2I-CompBench++ and T2V-CompBench, NarrLV, and VBench 2.0, with published benchmark reference rows included for context.
In human evaluation PreviewDiff has the highest three-way preference and pairwise win rates against EvoSearch and Best-of-N, with larger gains on content than on style. The critic is prompted to focus on prompt adherence, so the search directly optimizes semantic correctness rather than aesthetics; weaker style gains are attributed to the current critic objective. Nine human raters, 360 image prompts from GenEval2 and 200 video prompts from NarrLV, with 9 independent ratings per prompt and criterion, hidden method names, randomized order, and bootstrapped confidence intervals for three-way preference.
Perspective
The method targets pretrained diffusion or flow-matching models that expose intermediate clean-latent estimates and support controlled local restarts, and it requires an accessible multimodal critic; the paper validates it on SDXL, SD-3.5, FLUX.2, LTX-Video, and Wan 2.2, with GenEval2 as the main image scaling benchmark and NarrLV as the main video scaling benchmark. It suits settings where prompt adherence matters more than aesthetics, since the critic is prompted to focus on semantic correctness such as objects, attributes, and counts. For teams that want to trade extra verifier compute for better compositional fidelity without finetuning a model, the pipeline offers a directly reusable combination of checkpoint decoding, semantic-correction branching, and beam pruning.
Search width gives the largest gains, while depth and semantic variants give smaller and saturating gains; earlier checkpoints help but very early clean-latent estimates can be unreliable, a tradeoff the paper balances with bounded checkpoints and small restart depths. Remaining failures cluster around exact counting, inverted spatial relations, attribute binding, and long-horizon temporal state tracking, such as prompts requiring four clocks and five spotted turtles, sheep and umbrellas in a specific vertical relation, a blue giraffe, or a transparent-to-blue-to-green-to-yellow-to-orange-to-purple color sequence. Style gains are weaker than content gains, which the authors attribute to the critic objective rather than model capability. The paper also notes that a restart-depth ablation would isolate the effect of the local restart itself, so readers may watch how that axis develops in follow-up work.
