It’s All About One Direction
Abstract
Diffusion Transformers exhibit several sparse structures, including high-norm tokens, massive activation channels, and attention sinks, but their relationship remains unclear. Across FLUX.1-Dev, FLUX.1-Schnell, and PixArt-, we find that high-norm register tokens collapse onto a direction shared across prompts and seeds. This direction is dominated by a single massive activation channel that accounts for most of the register tokens' high norm. Controlled interventions show that direction, rather than magnitude, determines sink identity, while large norms stabilize the register direction across subsequent layers. Query–key geometry explains this directional preference. The -aligned state primarily affects image detail, and its effect depends on the Transformer layer at which it is present. Finally, we trace a coordinated lifecycle in which projection onto grows as high-norm registers and attention sinks emerge, persists through intermediate layers, and is later attenuated as these sparse structures disappear. Together, these results link massive activation channels, high-norm tokens, attention sinks, and image refinement in Diffusion Transformers through a shared internal direction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.