Compensating Token Merging for Efficient Diffusion Transformers
Abstract
Token merging approaches accelerate diffusion transformers (DiTs) by aggregating redundant tokens into representative ones to reduce the quadratic cost of attention. Recent approaches further employ advanced selection criteria to identify redundant tokens for aggregation. Despite their effectiveness, they still suffer from significant merging errors, the discrepancies between full-token and merged-token outputs, leading to blurred textures and local artifacts. However, directly computing these errors requires attention over the full tokens, which negates the efficiency gain of token merging. To address this, we propose ComE-ToMe, a novel framework that compensates for the merging errors efficiently. Specifically, we introduce calibrated window attention (CWA) that estimates the errors within local neighborhoods to restore the lost details. By computing attention within local windows while incorporating outside-window tokens through their averaged keys, CWA effectively captures fine details without the cost of full attention. We further introduce compensation-aware merging (CAM), which adapts token merging for CWA by preserving boundary tokens and using cumulative CFG importance to stably protect salient regions across denoising steps. Extensive experiments across various DiT architectures demonstrate that ComE-ToMe consistently outperforms prior approaches while maintaining efficiency. We will make our code publicly available online upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.