Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

FADE erases 16 concepts at once from a single text-to-video model, cutting residual object accuracy to 4.9% from the strongest baseline's 15.5% while staying within 0.9% of the unedited VBench average

The work proposes FADE, a multi-concept unlearning framework for text-to-video diffusion transformers that first applies a joint closed-form key/value edit suppressing all target concepts, then trains per-concept frame-aware low-rank adapters gated by the frame index and denoising timestep to remove residual per-frame leakage, and combines adapters by a similarity-based soft router; erasing 16 concepts (objects, artistic styles, and nudity) from a single Wan2.1-T2V-1.3B backbone, FADE reduces residual object-benchmark accuracy to 4.9% against 15.5% for the strongest of eight baselines, keeps the VBench average within 0.9% of the unedited model, and the ranking holds under a VLM judge and a blinded human study.