acceptodds
Under review as a conference paper at ICLR 2027

Beyond Positional Adaptation: Calibrating Attention for Image Resolution Extrapolation

Abstract

Diffusion transformers have rapidly advanced image generation, yet training them at higher resolutions remains expensive because dense attention scales quadratically with the number of image tokens. A prominent approach to training-free resolution extrapolation adapts rotary positional embeddings to spatial coordinates beyond the training range. However, higher resolution also expands the pool of tokens competing within each softmax, which can dilute attention to informative tokens even after positional adaptation. We introduce Cardinality-Calibrated Attention (CCA), a plug-and-play correction that compensates for this change in token count. CCA rescales attention logits by , where and are the current and native-reference key counts of each attention pool. The rule retains the chosen positional mapping and requires neither retraining nor fitted calibration parameters. It exactly recovers the native attention operator at the reference key count. Experiments on FLUX and Qwen-Image across multiple benchmarks demonstrate that CCA improves image quality in training-free resolution extrapolation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.