Public articles linked to the same research event.
arXiv The work introduces Pivot-SD, an offline self-distillation framework that selects high-impact commitments (pivots) during denoising via an information-gain metric measuring uncertainty reduction over remaining masked positions, trains pivots from successful trajectories with cross-entropy and pivots from failed trajectories with targeted unlikelihood while leaving the rest of the failed trajectory untouched, and with only 200 questions and four rollouts each improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
The work introduces Pivot-SD, an offline self-distillation framework that selects high-impact commitments (pivots) during denoising via an information-gain metric measuring uncertainty reduction over remaining masked positions, trains pivots from successful trajectories with cross-entropy and pivots from failed trajectories with targeted unlikelihood while leaving the rest of the failed trajectory untouched, and with only 200 questions and four rollouts each improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
The work introduces Pivot-SD, an offline self-distillation framework that selects high-impact commitments (pivots) during denoising via an information-gain metric measuring uncertainty reduction over remaining masked positions, trains pivots from successful trajectories with cross-entropy and pivots from failed trajectories with targeted unlikelihood while leaving the rest of the failed trajectory untouched, and with only 200 questions and four rollouts each improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
The work introduces Pivot-SD, an offline self-distillation framework that selects high-impact commitments (pivots) during denoising via an information-gain metric measuring uncertainty reduction over remaining masked positions, trains pivots from successful trajectories with cross-entropy and pivots from failed trajectories with targeted unlikelihood while leaving the rest of the failed trajectory untouched, and with only 200 questions and four rollouts each improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.