Skip to main content
Back to timeline
arXivSource publication:

AnchorCache pairs static text anchors with attention-mask co-design, letting in-context diffusion generation match full-attention quality across image, speech, and video while reaching up to 6.40x inference speedup

Synopsis

The work introduces AnchorCache, a parameter-free token-layout and attention-mask co-design that inserts static text anchors so reference representations are conditioned on the instruction during cache construction, after which reference keys and values can be reused exactly across denoising steps; to recover the quality initially lost through this structural conversion, the authors apply teacher-forced velocity distillation followed by a short on-policy stage that queries the teacher at student-visited states, matching full-attention quality across image, speech, and video generation benchmarks with efficiency gains that grow with reference-context size and reach a 6.40x speedup in diffusion transformer inference.

Source-provided article image: Beyond Attention Masks: Instruction Anchoring for Efficient In-Context Diffusion Generation
Figure 2 · arXiv

Interpretation

AnchorCache inserts static text anchors so that reference representations are conditioned on the instruction during cache construction, after which reference keys and values can be reused exactly across denoising steps. Previously, decoupling reference tokens from the target enabled exact key-value reuse but prevented references from attending to the instruction, degrading instruction following and reference fidelity; the authors state this trade-off cannot be resolved through attention-mask design alone, and AnchorCache sidesteps it through token-layout and mask co-design. The abstract describes the mechanism at the method level and reports quality matching full attention across benchmarks spanning image, speech, and video generation; specific benchmark names, model scales, and per-benchmark metrics are not listed in the abstract.

To recover the quality initially lost through the structural conversion, the authors apply teacher-forced velocity distillation followed by a short on-policy stage that queries the teacher at student-visited states. The authors state this is the first use of on-policy distillation for architectural recovery in diffusion models. Presented in the abstract as a method description plus an author-stated first-of-its-kind claim, without ablation numbers or training-cost details for the distillation stages.

Efficiency gains increase with reference-context size, reaching a 6.40x speedup in diffusion transformer inference. The gain comes from no longer repeating reference-side computation at every denoising step, and it amplifies as more references are added, which is exactly the part of in-context diffusion whose cost grows fastest. The abstract reports the 6.40x peak speedup figure and states the gain grows with reference-context size; a full speedup curve across reference sizes is not given.

Perspective

The result targets inference in in-context diffusion transformers that concatenate instruction, target, and reference tokens for joint attention, and it suits generation tasks where reference context is large and repeated reference-side computation dominates cost; beneficiaries are image, speech, and video applications and inference services that need multi-reference conditional generation. The method is described as parameter-free, so its applicability presupposes a model that supports the designed token layout and attention mask.

The abstract does not list specific benchmark names, model scales, reference-context sizes, or per-item quality metrics, so the boundary conditions of 'matches full-attention quality' still need checking in the main text; 6.40x is the reported peak speedup, and the speedup curve across reference sizes and the quality-efficiency trade-off point remain unclear; the relative contribution of teacher-forced velocity distillation versus the short on-policy stage, along with extra training cost and stability, are also questions a reader would keep watching. The text available here is abstract and metadata level, with figures and experiment tables not included, which limits further summarization of those details.

Sources