acceptodds
Under review as a conference paper at ICLR 2027

Spatially Adaptive Multi-Resolution Token Merging for Efficient Diffusion Inference

Abstract

Diffusion transformers repeatedly process dense spatial token grids, making high-resolution generation computationally expensive. Existing training-free token-merging methods primarily use global token matching or ranking, overlooking that redundancy in generated images is spatially structured. We formulate diffusion token reduction as budgeted hierarchical spatial-resolution allocation and introduce **QuadMerge**, which converts classifier-free-guidance residuals into a multi-resolution partition of , , and regions. QuadMerge conservatively coarsens only uniformly low-importance regions, aggregates them with importance-weighted pooling, couples the plan with timestep-adaptive compression, and caches allocation plans across compatible transformer blocks. At a matched token budget, QuadMerge improves both generation quality and measured latency over global token-merging baselines, and the margin grows under aggressive compression. On PixArt-α, a 30% token reduction costs 0.09 FID (36.60 → 36.69) versus 1.05 for importance matching and 12.48 for ToMe; unlike those matchers, it stays below uncompressed wall-clock cost. At a 70% reduction the FID gap over importance matching grows to 7.92, and the same allocation transfers to Stable Diffusion 2.1. The method requires no retraining, no additional learned modules, and no modification to the diffusion sampler.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.