Public articles linked to the same research event.
arXiv The work introduces a post-training method for diffusion language models that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the contextual token-feature space of a frozen pretrained DLM, optimizing discrete models with REINFORCE plus a leave-one-out baseline and continuous models by differentiating through generated latents, so that no full sampling trajectories or jointly trained auxiliary models are needed; it reports lower generative perplexity at comparable entropy on OpenWebText, better accuracy-computation trade-offs on GSM8K, and increased decoding parallelism at similar or higher accuracy on 16B DMax-LLaDA2.0 hybrid masked-uniform diffusion models.
The work introduces a post-training method for diffusion language models that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the contextual token-feature space of a frozen pretrained DLM, optimizing discrete models with REINFORCE plus a leave-one-out baseline and continuous models by differentiating through generated latents, so that no full sampling trajectories or jointly trained auxiliary models are needed; it reports lower generative perplexity at comparable entropy on OpenWebText, better accuracy-computation trade-offs on GSM8K, and increased decoding parallelism at similar or higher accuracy on 16B DMax-LLaDA2.0 hybrid masked-uniform diffusion models.
The work introduces a post-training method for diffusion language models that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the contextual token-feature space of a frozen pretrained DLM, optimizing discrete models with REINFORCE plus a leave-one-out baseline and continuous models by differentiating through generated latents, so that no full sampling trajectories or jointly trained auxiliary models are needed; it reports lower generative perplexity at comparable entropy on OpenWebText, better accuracy-computation trade-offs on GSM8K, and increased decoding parallelism at similar or higher accuracy on 16B DMax-LLaDA2.0 hybrid masked-uniform diffusion models.
The work introduces a post-training method for diffusion language models that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the contextual token-feature space of a frozen pretrained DLM, optimizing discrete models with REINFORCE plus a leave-one-out baseline and continuous models by differentiating through generated latents, so that no full sampling trajectories or jointly trained auxiliary models are needed; it reports lower generative perplexity at comparable entropy on OpenWebText, better accuracy-computation trade-offs on GSM8K, and increased decoding parallelism at similar or higher accuracy on 16B DMax-LLaDA2.0 hybrid masked-uniform diffusion models.