Skip to main content
Back to timeline
arXivSource publication:

EVA replaces the standard Gaussian latent prior with an empirical autoregressive prior, matching competitive generation quality on ImageNet and VGGSound with fewer inference parameters

Related research and updates

Synopsis

The authors introduce Empirical Variational Autoencoder (EVA), which keeps the VAE evidence lower bound but replaces the fixed standard-Gaussian latent prior with an autoregressive prior learned empirically from data, aligning prior with posterior; on a controlled toy benchmark, ImageNet 256×256 image generation, and VGGSound sound generation, EVA achieves competitive results against autoregressive generators with GMM or diffusion heads while ancestral inference is substantially faster than diffusion-based autoregressive approaches, and under matched inference-time Transformer depth it attains lower FID than AR-Diffusion with fewer inference parameters.

AI-generated editorial illustration: Empirical Variational Autoencoder

Interpretation

EVA replaces the VAE's fixed standard-Gaussian latent prior with an autoregressive prior learned empirically from training data by an additional predictor network, conditioning each latent token on all preceding latent positions so the prior can represent dependencies across the entire preceding latent sequence. Conventional VAEs use a standard Gaussian prior shared across spatial or temporal indices, which mismatches residual dependencies in the posterior; prior work strengthens priors via VampPrior, resampled priors, flow-based priors, or hierarchical VAEs, but hierarchical VAEs still use factorized conditional distributions within each spatial latent map. EVA's change is a single linear layer on top of Causal VAE to predict the latent prior. The paper derives the EVA ELBO (Eqs. 15–17) and its conditional-generation extension (Eqs. 18–23), and compares four generative families theoretically in a toy experiment.

On the controlled toy benchmark, EVA places fewer samples between true modes and concentrates mass more tightly on the true modes. MSE-trained AR falls into the averaging trap and fails to recover the multimodal joint structure; the GMM head is limited by a pre-defined number of mixture components and is forced to cover multiple modes with a single component when visible modes exceed capacity; Causal VAE captures the major clusters but still generates a noticeable amount of inter-mode samples. The toy process uses a one-dimensional latent sequence and one-dimensional observations, with latent transition modes set to 1, 2, or 3 and observation modes set to 1, 2, or 3; results are shown as samples in a two-dimensional space (Fig. 3).

On ImageNet 256×256 and VGGSound, EVA achieves competitive FID at comparable inference cost and is substantially faster than diffusion-based autoregressive methods. AR-Diffusion (321M) achieves the best metrics but requires iterative denoising and is about 10 times slower at inference than EVA; under matched inference-time Transformer depth, AR-Diffusion (236M) remains at much worse FID, while EVA achieves substantially lower FID with fewer inference parameters. Causal VAE performs poorly, showing that merely making a VAE decoder causal is not sufficient for high-quality sequential generation. All models share a 170M-parameter, 24-block Transformer backbone with hidden width 768; 50k samples are generated for ImageNet and 15446 for VGGSound, with FID, IS, training and inference parameter counts, and normalized inference time reported (Table 1).

EVA improves generation FID while maintaining reconstruction fidelity comparable to Causal VAE, and develops a semantically structured latent space. Classical VAEs with a single global latent bottleneck trade off reconstruction fidelity against prior regularization; in the multi-latent setting, Causal VAE and EVA have similar rFID while EVA substantially improves generation FID, suggesting the two objectives need not be tightly coupled when sufficient latent capacity is available. Table 2 reports rFID comparisons; Fig. 4 uses bi-directional KL divergence to show that semantically similar patches have similar latent distributions in EVA's latent space, whereas Causal VAE's posterior latent distributions still exhibit spatially structured variations.

Perspective

The work targets generative modeling of continuous-valued sequences and applies to data that can be encoded into continuous latent tokens by a pretrained VAE, such as images and audio; it is validated on two closed-set benchmarks, ImageNet 256×256 and VGGSound, at moderate model scale (170M-parameter backbone, plus EVA-S and EVA-L variants), trained on eight NVIDIA RTX5090 (32GB) GPUs for about 3 days on ImageNet and 12 hours on VGGSound. Conditional generation is implemented via a conditional encoder, conditional decoder, and conditional predictor, and CFG can be applied to the mean of the predicted latent distribution. Latent-size analysis shows 128 dimensions gives the best FID on ImageNet, while excessively large latent spaces make prior modeling and sampling harder. The authors note next steps toward larger-scale and more diverse downstream tasks such as text-to-image, text-to-speech, and multimodal data generation.

EVA currently linearizes spatial latent maps into a 1D causal sequence and does not explicitly exploit 2D spatial structure; whether autoregressive priors specialized for spatial data further improve image generation remains an open question. Experiments are limited to closed-set generation benchmarks on ImageNet and VGGSound at moderate model scale, and behavior at larger scale and on more diverse tasks remains to be validated. The VGGSound dataset is currently provided only by third parties and lacks part of the original data, which may affect the reproducible scope of results on that benchmark. In addition, the latent-size and model-size analyses are based on CFG-scale sweeps on ImageNet, and optimal configurations for other modalities and data scales still need further examination.

Sources