Beyond Layers: Position-Resolved Gradient Conflict and Position-Aware Modulation for Unified Multimodal Models
Abstract
Unified multimodal models (UMMs) train image understanding and autoregressive image generation on shared parameters, and the two objectives are known to interfere. Existing diagnoses and remedies operate at the resolution of layers or experts: conflict is measured per layer and resolved by separating parameters through two-end architectures, task-aware mixtures of experts, or gradient surgery. We argue that this resolution hides an orthogonal axis. Generation in a UMM is next-token prediction over a raster sequence of visual tokens whose roles vary systematically with position—early tokens fix global layout and low-frequency structure, late tokens fill in texture—so how strongly a generation gradient interferes with understanding should depend on where in the sequence it originates. We introduce a position-resolved interference map that attributes the understanding–generation gradient conflict to visual-token positions within every layer, computed from a single backward pass with position-grouped gradient hooks at the cost of a standard backward pass. On Show-o and Janus-Pro, position explains a substantial share of conflict variance after controlling for depth (partial vs. for layer on Show-o; vs. on Janus-Pro, whose understanding branch bypasses the visual-token sequence): the first quarter of the sequence has a mean gradient cosine of against understanding, the last quarter . The dependence survives per-position gradient-norm normalization, retaining of its effect size, so it is directional rather than a magnitude artifact, and conflict strength tracks the semantic content of a position (Spearman ). Building on the map, we propose position-aware modulation (PAM), which removes the anti-aligned component of generation gradients only at high-conflict positions without changing the architecture. Under a matched trainable-parameter budget PAM improves over layer-wise separation by MME and GenEval points while matching it on POPE and overall FID, whereas a random-position control with the same number of modulated tokens recovers only about of the gain; the cost is a relative degradation of the high-frequency component of FID, with overall FID unchanged. Position-based and layer-based separation are complementary degrees of freedom and can be combined.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.