acceptodds
Under review as a conference paper at ICLR 2027

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model

Abstract

Unified multimodal models (UMMs) integrate image understanding and generation within a single architecture, yet the two tasks require different visual representations, creating inherent conflicts during joint training. Recent methods increasingly improve UMMs through architectural decoupling (e.g., decoupled visual encoders, MoE/MoT architectures, or frozen MLLMs), but why decoupling works remains poorly understood. In this work, we revisit architectural decoupling through layer-wise cross-modal attention interaction patterns. Our analysis reveals a previously overlooked phenomenon: as architectural decoupling becomes stronger, understanding and generation retain distinct interaction regimes, while each task's interaction pattern becomes increasingly similar to that of a strong task-specific reference model. This suggests that the benefit of architectural decoupling may stem from enabling **task-specific cross-modal interaction specialization**, rather than driving the two tasks toward a shared interaction behavior. Motivated by this finding, we propose **A**ttention **I**nteraction **A**lignment (**AIA**), which adaptively transfers task-specific reference interaction behavior without further architectural decoupling. Specifically, **stage-level interaction alignment** transfers the overall cross-layer interaction knowledge of the reference while preserving the student's own representation organization, whereas a **Huber loss** retains the model's ability to adaptively adjust interaction across individual layers. Applied to Emu3 during supervised fine-tuning and Janus-Pro during post-training, AIA consistently improves both understanding and generation, and further improves interleaved cross-task performance. These findings suggest that strong unified multimodal performance does not require understanding and generation to converge to a shared interaction regime; instead, each task can retain and exploit its own appropriate interaction behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.