acceptodds
Under review as a conference paper at ICLR 2027

Beyond Attention Masks: Instruction Anchoring for Efficient In-Context Diffusion Generation

Abstract

In-context diffusion transformers concatenate instruction, target, and reference tokens into a single sequence for joint attention. Reference-side computation must be repeated at every denoising step, with computation growing rapidly as more references are added. Decoupling the reference tokens from the target enables exact K/V reuse across denoising steps. This isolation, however, prevents the reference tokens from attending to the instruction, degrading instruction following and reference fidelity. This trade-off cannot be resolved by attention-mask design alone. We introduce AnchorCache, a parameter-free token-layout and attention-mask co-design that inserts static text anchors. The anchors condition the reference representations on the instruction during cache construction, after which the resulting reference K/V are reused exactly across denoising steps. This structural conversion initially degrades generation quality. We recover the lost performance with teacher-forced velocity distillation, followed by a short on-policy stage that queries the teacher at student-visited states. To our knowledge, this is the first use of on-policy distillation for architectural recovery in diffusion models. Across benchmarks spanning image, speech, and video generation, AnchorCache matches full-attention quality. Its efficiency gains increase with the reference context, reaching a 6.40× speedup in DiT inference. Code is available in an anonymous repository: https://anonymous.4open.science/r/AnchorCache-FD70/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.