Public articles linked to the same research event.
arXiv The authors introduce Empirical Variational Autoencoder (EVA), which keeps the VAE evidence lower bound but replaces the fixed standard-Gaussian latent prior with an autoregressive prior learned empirically from data, aligning prior with posterior; on a controlled toy benchmark, ImageNet 256×256 image generation, and VGGSound sound generation, EVA achieves competitive results against autoregressive generators with GMM or diffusion heads while ancestral inference is substantially faster than diffusion-based autoregressive approaches, and under matched inference-time Transformer depth it attains lower FID than AR-Diffusion with fewer inference parameters.
The authors introduce Empirical Variational Autoencoder (EVA), which keeps the VAE evidence lower bound but replaces the fixed standard-Gaussian latent prior with an autoregressive prior learned empirically from data, aligning prior with posterior; on a controlled toy benchmark, ImageNet 256×256 image generation, and VGGSound sound generation, EVA achieves competitive results against autoregressive generators with GMM or diffusion heads while ancestral inference is substantially faster than diffusion-based autoregressive approaches, and under matched inference-time Transformer depth it attains lower FID than AR-Diffusion with fewer inference parameters.
The authors introduce Empirical Variational Autoencoder (EVA), which keeps the VAE evidence lower bound but replaces the fixed standard-Gaussian latent prior with an autoregressive prior learned empirically from data, aligning prior with posterior; on a controlled toy benchmark, ImageNet 256×256 image generation, and VGGSound sound generation, EVA achieves competitive results against autoregressive generators with GMM or diffusion heads while ancestral inference is substantially faster than diffusion-based autoregressive approaches, and under matched inference-time Transformer depth it attains lower FID than AR-Diffusion with fewer inference parameters.
The authors introduce Empirical Variational Autoencoder (EVA), which keeps the VAE evidence lower bound but replaces the fixed standard-Gaussian latent prior with an autoregressive prior learned empirically from data, aligning prior with posterior; on a controlled toy benchmark, ImageNet 256×256 image generation, and VGGSound sound generation, EVA achieves competitive results against autoregressive generators with GMM or diffusion heads while ancestral inference is substantially faster than diffusion-based autoregressive approaches, and under matched inference-time Transformer depth it attains lower FID than AR-Diffusion with fewer inference parameters.