Prompt Reinjection for Spatial Reasoning in MMDiTs: Signal Amount, Depth Allocation, and Content-Specific Repair
Abstract
Multimodal diffusion transformers (MMDiTs) forget prompt semantics as depth grows, and spatial relations suffer most. Prompt Reinjection (PR) counters this by repairing every text token at every layer, but neither the number of repaired positions nor the choice of layers has been examined; we therefore treat PR as a probe of its own redundancy. Repair saturates early: eight real positions match full-scale repair, and the injection weight, not the token count, controls the response; past a bounded window, amplification collapses performance on SD3. Most repaired positions are padding, which joint attention does not mask and which carries no signal. Depth adds a secondary trend at the default weight: with span-matched controls, dense repair over the early half matches full PR while late-only repair is no better than baseline, and several allocations perform comparably. As one instantiation, an early-full/late-sparse schedule (Layer-Adaptive Sparse Prompt Reinjection, LASP) reaches 40.1% on GenEval Position, matches full PR with a third of the token-layer pairs, and remains robust where full PR collapses; we make no optimality claim. The injected content matters—noise, a permutation, or a constant removes the gain—and the pattern transfers to T2I-CompBench and directionally to FLUX.1-Schnell.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.