acceptodds
Under review as a conference paper at ICLR 2027

Probing Cross-Branch Compatibility in Unified Multimodal Models through Semantic Steering

Abstract

Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet these two capabilities are still largely evaluated in isolation. Strong performance on both tasks therefore tells us little about how their internal semantic representations relate to each other. A fundamental question is whether the semantic structures underlying understanding and generation are compatible across branches. To study this question, we introduce cross-branch semantic steering, an intervention-based framework that extracts semantic directions from one branch and applies them to the other. % We show that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness. In contrast, the reverse direction consistently shows limited effectiveness. Our analysis suggests that this asymmetry may be related to a practical representational mismatch: understanding-derived vectors capture transferable, object-centric semantics, while generation-derived vectors primarily encode low-level appearance features. Our results reveal that architectural unification does not guarantee semantic alignment, and establish cross-branch steering as a practical tool for probing multimodal representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.