Commutator-Guided Calibration of Attention-Head Interactions in Diffusion Transformers
Abstract
Diffusion Transformers (DiTs) aggregate attention-head outputs additively without explicitly composing their state-dependent residual operators. We diagnose their non-commutativity (NC), or dependence on composition order, using routing-matrix commutators and changes in predicted clean latents under reversed head-operator composition. Their calibrated comparison defines weighted undershoot graphs: Laplacian traces summarize layer-level mismatch, while edge weights identify pairs with weak functional responses relative to their routing-level reference. We use a two-phase procedure: Phase 1 probes a frozen pretrained model to select compact, layer-specific head-pair sets without generation metrics; Phase 2 fine-tunes the model with a one-sided hinge objective applied to these pairs. Across ImageNet and CUB-200-2011, the results suggest that intervention efficacy reflects both calibrated undershoot and layer position: high diagnostic energy alone does not guarantee larger generation gains. Across continuation seeds, simultaneous NC regularization at layers 10 and 24 achieves ImageNet validation FID , compared with for pretrained DiT-XL/2, and improves Inception Score, and precision over matched denoising-only (no-NC) fine-tuning. On CUB-200-2011, NC achieves FID , compared with for the domain-adapted baseline, with higher precision. No-NC controls support gains beyond continued denoising training, while pair-budget ablations favour compact interventions over broader regularization. These results support efficacy of selective regularization of attention-head interactions while retaining the additive DiT architecture.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.