acceptodds
Under review as a conference paper at ICLR 2027

When Co-Training Becomes Redundant: Update Geometry in Unified Multimodal Reinforcement Learning

Abstract

Unified multimodal models (UMMs) combine multimodal understanding and generation in one backbone through separate experts. In image-generation reinforcement learning (RL), the reward signal backpropagates through both experts, so it can update the understanding expert as well as the generation expert, yet the role of the former's updates remains unclear as prior work typically keeps it frozen. In this work, we show that the understanding expert's contribution is front-loaded. Two velocity-space probes including compensability (measuring whether its updates carry content the generator cannot locally reproduce) and transport (measuring how strongly its updates align with directions the generator realizes efficiently), reveal that the early updates of the understanding expert are both distinct and efficient, while both signals decay as training proceeds. This finding is also mirrored at the reward level, where later understanding training steps add no measurable gain. These observations motivate a staged protocol: a geometry-based freeze rule that reads the two probes relative to their running peak, and freezes the understanding expert at a predicted step. Across two backbones, three reward models, and three RL algorithms, staged training consistently outperforms naive joint training, improving GenEval2 by up to while reaching higher rewards in substantially fewer training hours. Our results suggest the understanding expert's updates are best treated as a trajectory-dependent, front-loaded resource rather than a fixed co-training partner.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.