acceptodds
Under review as a conference paper at ICLR 2027

Compensating Token Merging for Efficient Diffusion Transformers

Abstract

Token merging approaches accelerate diffusion transformers (DiTs) by aggregating redundant tokens into representative ones to reduce the quadratic cost of attention. Recent approaches further employ advanced selection criteria to identify redundant tokens for aggregation. Despite their effectiveness, they still suffer from significant merging errors, the discrepancies between full-token and merged-token outputs, leading to blurred textures and local artifacts. However, directly computing these errors requires attention over the full tokens, which negates the efficiency gain of token merging. To address this, we propose ComE-ToMe, a novel framework that compensates for the merging errors efficiently. Specifically, we introduce calibrated window attention (CWA) that estimates the errors within local neighborhoods to restore the lost details. By computing attention within local windows while incorporating outside-window tokens through their averaged keys, CWA effectively captures fine details without the cost of full attention. We further introduce compensation-aware merging (CAM), which adapts token merging for CWA by preserving boundary tokens and using cumulative CFG importance to stably protect salient regions across denoising steps. Extensive experiments across various DiT architectures demonstrate that ComE-ToMe consistently outperforms prior approaches while maintaining efficiency. We will make our code publicly available online upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.