FADE erases 16 concepts at once from a single text-to-video model, cutting residual object accuracy to 4.9% from the strongest baseline's 15.5% while staying within 0.9% of the unedited VBench average
Related research and updatesSynopsis
The work proposes FADE, a multi-concept unlearning framework for text-to-video diffusion transformers that first applies a joint closed-form key/value edit suppressing all target concepts, then trains per-concept frame-aware low-rank adapters gated by the frame index and denoising timestep to remove residual per-frame leakage, and combines adapters by a similarity-based soft router; erasing 16 concepts (objects, artistic styles, and nudity) from a single Wan2.1-T2V-1.3B backbone, FADE reduces residual object-benchmark accuracy to 4.9% against 15.5% for the strongest of eight baselines, keeps the VBench average within 0.9% of the unedited model, and the ranking holds under a VLM judge and a blinded human study.
Figure 1 : Frame reactivation and multi-concept erasure. Prior T2V erasure methods apply the same suppression to every frame, and an erased concept can resurface mid-clip (red boxes; top: Refusal Vector on parachute , middle: VideoEraser on English springer ). FADE erases several concepts from a single backbone, and in these examples they stay suppressed in every frame (bottom).
arXivInterpretation
The paper formalizes the frame-reactivation gap, the difference between peak-frame and clip-mean concept presence, and shows it exposes failures that clip-level accuracy misses: removing frame conditioning from FADE leaves clip-mean accuracy almost unchanged but triples the gap. Prior T2V erasure evaluations rely mainly on clip-level averages, which average away concepts that resurface in only a few frames; this diagnostic adds per-frame peak presence to the metric. The ablation (Table 8) shows that removing FrameGate changes accuracy only from 4.9 to 4.2 while raising the reactivation gap from 2.7 to 8.5 and FVD by 23.8.
FADE combines a joint closed-form K/V edit (Concept Redirect, with motion-paraphrase embeddings and texture-phase decay), per-concept FA-LoRA adapters gated by frame and timestep, and similarity-based C-MoE routing, which the authors describe as the first framework to erase multiple concepts from a single T2V model. Existing T2V methods apply the same suppression to every frame and are evaluated with one target concept or category at a time; FADE adds frame- and timestep-dependent suppression and multi-concept composition. The method shares all FAE hyperparameters across four backbones (Wan2.1-1.3B/14B, CogVideoX-2B, HunyuanVideo-1.5-480P), recalibrating only router thresholds, and removes 87-94% of concept presence.
With 16 concepts, FADE lowers residual object accuracy from 77.8 to 4.9 (a 93.7% relative reduction), is the strongest eraser on each of nine classes, and raises FVD by only 7.8% against 25.5-46.5% for the baselines. The strongest baseline, VideoEraser, reaches 15.5, while the other baselines leave 52.6-65.7; directly ported T2I methods (UCE, ESD-x, ESD-u, MACE) remove less than a third of concept presence. All methods share the backbone, prompts, seeds, sampler, and judges; the ranking is the same under ResNet-50, a VLM (Qwen3.6-27B), and a blinded human study (FADE at 4.9/4.8/5 versus VideoEraser at 15.5/14.7/15).
FADE keeps its lead on compositional prompts, on 30 simultaneously erased celebrity identities, and under four jailbreak suites: 4.0 residual on compositional prompts (VideoEraser 12.9), 6.4 recognition on erased identities with 86.1 retained, and the lowest attack success rate across all four suites. Prior T2V methods did not evaluate prompts combining several erased concepts, large identity sets, or cross-suite adversarial robustness. Compositional values are close to single-target prompts on the same model, though the authors note the prompts differ, so this is not a controlled measurement of interference; the celebrity experiment is the largest concept set (30 identities); under jailbreaks the closest competitor, VideoEraser, has 1.3-1.7 times FADE's attack success rate.
Perspective
The result targets operators of deployed open-weight text-to-video diffusion transformers who want to remove explicit, violent, hateful, or copyrighted content and specific people's likenesses without full-model fine-tuning. It applies to erasing several named concepts at once (objects, artistic styles, nudity, celebrity identities), is evaluated on clips up to 81 frames, and is trained and tested on four backbones including Wan2.1-T2V-1.3B. The method requires concept names and per-concept training: the joint CR solve takes about one minute, each expert trains in about 4 hours and stores 1.15 MB, 16 concepts total 62 GPU-hours, and experts train independently so runs parallelize.
No attack designed against FADE itself is evaluated, such as prompts that keep a concept's router gate below threshold or shift the concept toward weakly gated frames and timesteps. The largest concept set has 30 identities, and expert separation and router discriminability may decrease as the concept count grows; very fine distinctions such as two periods of one artist, and experts with opposed updates on overlapping inputs, lie outside the experiments. Modifier tokens such as English or van can raise an unrelated gate at some fidelity cost. Adding a concept after deployment would require re-solving CR and would leave the new concept out of existing experts' hard negatives, an incremental setting not evaluated. When the backbone cannot generate a concept (tench), calibration disables its expert. On evaluation, automated judges score frames and humans rate clips, while video-native judges of motion semantics are not yet established; clips beyond 81 frames and autoregressive generators are not evaluated.
