Adversarial post-training restores missing high-frequency detail in pixel diffusion, cutting DeCo FID from 33.27 to 28.59
Synopsis
By keeping the original diffusion or flow-matching objective and adding an adversarial loss to the predicted output at non-high-noise timesteps, this study jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality on the DeCo and PixelGen pixel backbones, attributing the gain to restored missing natural-image high-frequency power, while the same procedure yields no comparable gain in the tested latent diffusion configurations.
Interpretation
Adversarial post-training delivers joint gains in pixel diffusion: DeCo improves FID from 33.27 to 28.59, recall from 0.361 to 0.406, and DPG Score from 81.6 to 83.3, with PixelGen improving on every axis as well. The authors describe this as the first systematic study of adversarial post-training for pixel diffusion, and it changes neither model architecture nor sampling procedure, uses no distillation, does not reduce sampling steps, and uses no preference labels or reward model. Two pixel backbones fine-tuned from public checkpoints on the same BLIP3o-60k data, with GAN and no-GAN controls sharing starting checkpoint, data, and training horizon, evaluated on DPG-Bench and COCO-30k across prompt alignment, distribution fidelity and diversity, and no-reference quality.
The gain traces to restored natural-image high-frequency statistics: DeCo's radial power-spectrum slope moves from 2.59 to 2.24, close to the natural power law of roughly 2.0 for real COCO images, with about 0.35 dex more high-band power. Frequency-band and power-law analyses locate the improvement as measurable spectral power restoration rather than generic sharpening, and a spectrum-matched unsharp-mask control recovers only a small fraction of the GAN's FID gain. Luminance channels are Hann-windowed, transformed with a 2-D FFT, azimuthally averaged over constant-frequency rings, averaged across images, and fit by least squares, alongside band-power shares and unnormalized high-band log-power changes.
The gain is not mode dropping or memorization: recall rises rather than falls, and DINOv2 nearest-neighbor similarity stays at 0.586 on average while the maximum drops from 0.943 to 0.929. Adversarial training usually costs diversity, but as a gated post-training term it adds high frequency while raising coverage, and nearest-neighbor, recall, and matched-step no-GAN SFT controls rule out memorization, mode dropping, and additional optimization as simple explanations. A frozen DINOv2 encoder computes cosine similarity between each generated image and all training images, retaining the maximum as the nearest-neighbor score, with mean and maximum reported over the evaluation set and all gains measured against matched-step no-GAN SFT controls.
Perceptual loss is an alternative sharpening route with different costs: it also raises high frequency and Laplacian sharpness but lowers saturation, contrast, and colorfulness and worsens FID, pFID, and DPG Score, while the same procedure yields no comparable joint gain in the tested latent diffusion configurations and adds almost no decoded high-frequency power. The matched perceptual-loss comparison separates adversarial correction from generic sharpening, and the pixel-latent contrast identifies direct output access to the image statistics being corrected as a key factor governing success. On DeCo the perceptual arm worsens FID to 34.14 and DPG Score to 81.4 versus 28.59 and 83.3 for the GAN; SANA and PixArt show only small high-band share changes, and a perturbation probe through the frozen PixArt VAE measures a weaker decoded-HF response.
Perspective
The result targets already-converged multi-step text-to-image pixel diffusion models as a post-training quality correction: the original diffusion or flow-matching objective is retained and an adversarial loss is added to the predicted output at non-high-noise timesteps, without changing architecture or sampling procedure. The authors validate it on two pixel backbones, DeCo and PixelGen, and report operating-point trends across discriminator design, noise gate, and adversarial weight for choosing a target metric balance. For teams using latent diffusion with a frozen decoder, the pixel-latent contrast indicates the same joint gains may not appear.
Discriminator design, noise gate, and adversarial weight are reported as empirical trends and operating points; the authors state they do not define a complete ordering or a universally optimal configuration for every metric. Only several configurations of PixArt and SANA were tested on the latent side, so the output-access explanation still awaits testing on more backbones. The three percentage-valued spectral diagnostics use different normalizations and evaluation sets and should not be compared numerically, and readers working from the abstract alone would need the full text for exact numbers and ablation details.
