Skip to main content
Back to timeline
arXivSource publication:

Swapping T5-small for T5Gemma-2 and distilling its decoded probabilities cut a continuous diffusion language model's generative perplexity to 17.8, beating GPT-2-M

Synopsis

Holding the ELF continuous diffusion language model fixed and changing only the text embedding space, this work finds that scaling embedding models along T5-small to T5Gemma-1 to T5Gemma-2 improves diffusibility (about 40% lower generative perplexity at the same entropy), but raw T5Gemma-2 embeddings are so discriminative that imperfect sampling lands on invalid embeddings; distilling the teacher decoder's probabilities as soft labels into a student encoder pulls alternative candidate embeddings closer, and the medium-sized model reaches Gen. PPL 17.8 (real text 15.4) on OpenWebText, outperforming GPT-2-M.

AI-generated editorial illustration: Scaling and Distilling Text Embeddings for Better Diffusibility

Interpretation

In a controlled experiment that fixes the ELF diffusion framework and replaces only the embedding model, larger embedding pretraining scale yields better diffusion generation: T5Gemma-1-B beats T5-small, T5Gemma-2 beats T5Gemma-1-B, and T5Gemma-2 is the only embedding space reaching real-text entropy, cutting generative perplexity by about 40% relative to T5-small at the same entropy. Earlier continuous diffusion language models either used older embedding models or trained embeddings jointly with the diffusion model, making the embedding space's contribution hard to isolate; fixing the diffusion side to ELF and changing only the embedding separates that contribution. The same ELF-B is trained on a 25% OpenWebText subset with 512-token sequences, comparing T5-small, T5-base, T5Gemma-1-S/B, T5Gemma-2-270M and ModernBERT via generative perplexity-entropy curves; the authors note encoders differ in pretraining data, so the controlled comparison is T5Gemma-2 versus the distilled student.

Raw T5Gemma-2 embeddings exhibit a failure mode in which they are hard to generate perfectly: a generated embedding can lie far from every candidate word and be counted invalid or uncertain; on average a generated sequence has uncertain positions, some of them invalid, where decoding confidence is low or the wrong word is produced. Prior work framed such phenomena as interpolation effects or mode averaging; here they are recast as diffusibility, defined as the ability to generate valid, decodable embeddings, with a quantifiable criterion based on candidate distance and decoder top-1 probability. Using the deployed decoding head's top-k candidates as reference, thresholds are set from the median nearest-neighbor distance between real embeddings and the top-1 probability distribution on real text; uncertain and invalid positions are counted over generated sequences and one sequence is shown in a 2D visualization, with the authors noting the exact threshold is not critical as long as it divides the teacher's bimodal distances.

A student encoder distilled from T5Gemma-2 decoder probabilities is more diffusible than the teacher: it pulls alternative candidate embeddings at a position closer, its generated embeddings sit nearer to candidates and return closer to the original sentence under round-trip denoising, and it improves both generative perplexity and MAUVE over the teacher under the same ELF-B/ELF-M. Compared with copying teacher embeddings via MSE or training with one-hot cross-entropy, only soft-label distillation produces a student that surpasses the teacher and reaches real-text entropy; the student is initialized from evenly spaced teacher layers and still beats the teacher at equal depth, so the gain is not from reduced capacity. Distillation runs on the full OWT with the student encoder minimizing KL to the teacher's renormalized top-k distribution; results on OWT-1024 and LM1B are reported as means and standard deviations over 3 seeds with multiple samples per point, alongside round-trip denoising and embedding-distance analyses.

On OpenWebText-1024, the distilled embeddings let ELF-M reach Gen. PPL 17.8 (real text 15.4) at entropy 5.45, better than GPT-2-M's 20.8 and raw T5Gemma-2's 19.3; on LM1B the student also beats raw T5Gemma-2 (50.5 versus 60.6). This provides a strong continuous diffusion language model baseline built from the embedding space, and shows continuous diffusion can be compared with discrete diffusion and autoregressive baselines under a shared evaluation protocol. All generated and real-text samples are re-tokenized with the GPT-2-Large tokenizer before computing generative perplexity and unigram entropy, with MAUVE using GPT-2-Large features; the sampling grid sweeps self-conditioning scale and NFE, selecting the lowest generative perplexity among points reaching real-text entropy, then evaluating over 3 seeds.

Perspective

The result targets research and engineering settings that use continuous diffusion language models: under a fixed ELF framework, frozen embedding models, OpenWebText and LM1B as the main evaluation corpora, and embedding models kept at modest scale, it shows that choosing a stronger embedding space and distilling it with soft labels improves generation quality and sampling stability. It can be used directly to select or reshape latent spaces for continuous diffusion language models, and offers a reproducible distillation path for adapting autoregressive foundation models into diffusion-ready embeddings; the distilled student halves encoding cost relative to T5Gemma-2, embeddings are frozen and cacheable during training, and sampling is affected only at the final decoding step, so the route has room to scale efficiently.

The authors list open questions: few-step and one-step generation (raw T5Gemma-2 under-integrates and even collapses at low NFE), practical-scale models for real-world tasks such as QA, and how to skip the step of adapting an autoregressive model into an embedding model. In addition, encoders differ in pretraining data, so cross-encoder rows are not fully controlled, and the authors note only T5Gemma-2 versus the distilled student is a controlled comparison; distillation costs some discriminative power (downstream classification such as SST-2 drops), finer-grained analysis of the embedding geometry is left for future work, and evaluation covers OpenWebText and LM1B, with LM1B treated by the authors as a problematic corpus serving as a check rather than a comparison.

Sources