FoMo turns the forking moment of a diffusion trajectory into automatic perceptual-distance labels, reaching 0.733 SROCC on PIPAL and beating metrics trained on human-annotated data
Synopsis
The work proposes FoMo: forward-noising a reference image to a sampled timestep and then denoising it independently, using that forking moment as a pointwise perceptual-distance label, so training data can be generated with no human annotation, and trains reference-based IQA metrics with a RankNet-style global ranking objective; across four benchmarks and seven backbones it achieves the best average performance, with LPIPS-Alex reaching 0.733 SROCC on PIPAL versus 0.577 for the same architecture trained on human-annotated KADID-10k labels.
Interpretation
It proposes a fully automated, annotation-free data generation pipeline: given a reference image, a forking step is sampled from the noise schedule, a noisy latent is constructed and denoised to yield a variant, and the forking step itself serves as the perceptual distance label for that pair. Reference-based IQA previously relied on MOS pointwise scores or 2AFC pairwise preference labels; MOS collection is expensive and noisy (the text notes KADID-10k required 30 crowdsourced ratings per image from over 2,200 subjects), while 2AFC captures only relative comparisons. FoMo replaces human annotation with diffusion generative dynamics and directly produces pointwise labels that support comparison between arbitrary image pairs. The paper reports 480k labeled pairs in training, 240k using real ImageNet references and 240k using FLUX-generated references; the appendix analyzes label variance for repeated generation at the same reference and forking step (120 ImageNet references, five forking steps, 4,800 images), finding small absolute standard deviations and coefficients of variation up to about 10%.
Two human studies verify that the forking moment agrees with human perceptual ordering: in the single-reference ranking study, 30 participants, 190 ImageNet validation references and 1,982 responses gave a Spearman rank correlation of 0.970 between forking-step ordering and human judgments (inter-rater correlation 0.960); in the cross-reference 2AFC study, 28 participants, 250 items and 5,713 responses showed individual agreement with the label, near-perfect Fleiss agreement, and majority-vote consensus matching the label on every item once the timestep gap exceeded 20. It converts the existing observation that diffusion trajectories fix coarse structure early and fine detail late into a testable perceptual-distance proxy, and provides cross-reference global-consistency evidence rather than only within-reference monotonicity. Both studies are controlled human experiments reporting participant counts, item counts and agreement statistics; additionally, four metrics already fitted to human judgments (LPIPS-Alex, LPIPS-VGG, DISTS, DreamSim) stand in for observers on 20,000 generated pairs, yielding SROCCs of 0.932, 0.915, 0.904 and 0.904, with over 99.4% agreement once the gap exceeds 20 steps.
Pointwise labels let the training objective move from triplet-wise pairwise comparison to within-batch global ranking: a ground-truth comparison matrix is built and a RankNet-style binary cross-entropy supervises the ordering across all pairs. 2AFC data only supports triplet binary classification and never exposes the global ranking during training; FoMo's pointwise labels support a global ranking objective, and ablations show RankBCE consistently beats 2AFC-style pairwise BCE and L1 regression, with FoMo remaining the better training source under that objective. Ablations on PIPAL with LPIPS-Alex and DINOv3 as representatives: LPIPS-Alex gives 0.733 for Rank, 0.592 for 2AFC and 0.440 for L1; DINOv3 gives 0.699 for Rank, 0.622 for 2AFC and 0.694 for L1. Batch-size analysis shows Transformer backbones improving with larger batches while the CNN backbone peaks at batch size 64.
Cross-generator and cross-backbone generalization: data generated with SD-1.5, SD-XL, SD3 and FLUX.1 each yields a working metric at a reduced 10k-pair budget, with FLUX.1 strongest for both backbones; timestep-range ablations show the first 35 of 50 steps matter most. It shows the method is not tied to a single diffusion model while indicating that generator quality affects results, and it points to a path of training a diffusion model on a new domain to extend the data. The cross-generator experiment draws five disjoint 10k subsets from a 50k pool per generator with a fixed training seed; on LPIPS-Alex, SD-1.5 gives 0.687, SD-XL 0.662, SD3 0.695 and FLUX.1 0.741, while on DINOv3 they give 0.574, 0.418, 0.517 and 0.672.
Perspective
The result targets reference-based image quality assessment, applying to restoration tasks such as super-resolution and denoising that need a distance between a restored image and a reference, and to evaluation settings that require a global ordering across image pairs. The method assumes access to a diffusion model for data generation; the default configuration shown is FLUX.1-dev with up to 50 sampling steps, forking steps sampled uniformly from 1 to 50, and 480k training pairs, trained for roughly 240k iterations on a single NVIDIA RTX A40. Beneficiaries include IQA researchers who need large-scale perceptual labels without an annotation budget, and builders of evaluation pipelines for image restoration, compression and generation. The paper also notes an extension path: when the target domain differs from the training distribution, one can train a diffusion model on that domain and use it to generate new samples, carrying the same pipeline into a new domain.
Worth watching: labels come from the diffusion model's generation process, so the metric may inherit that model's perceptual preferences; the paper also notes weaker performance on distortion families absent from the training data (additive pixel noise, corruption confined to a small region, and global photometric changes such as contrast change and mean shift), and the per-distortion-type analysis shows the CNN backbone trailing the strongest human-annotated baseline on a majority of families in the three legacy synthetic-distortion benchmarks. The forking construction is stochastic, so variants generated from the same reference at the same forking step are not identical; the paper argues via a 4,800-image variance analysis that this perturbation is far smaller than the label separation, but that conclusion depends on the reference set and forking-step range measured. The human studies are also limited in scale (30 participants in the single-reference study, 28 in the cross-reference study), and in the narrowest timestep-gap bin human participants themselves agree least, indicating limited discriminative power near ties. The paper reports rank-correlation measures such as SROCC and does not present absolute-error calibration conclusions across datasets in the main text.
