acceptodds
Under review as a conference paper at ICLR 2027

Cross-Feature Guided Reinforcement Learning for Unified Multimodal Understanding and Generation

Abstract

Unified multimodal models (UMMs) integrate multimodal understanding and image generation, but joint reinforcement learning (RL) often suffers from optimization conflict because the two branches are optimized together with limited interaction. To address this issue, we propose **Cross-Feature Guided Reinforcement Learning (CGRL)**, which enables reciprocal guidance between the two branches during joint RL. Leveraging the semantic correspondence between paired images and captions, CGRL introduces cross-feature guidance by comparing each branch's source representation with the generated token hidden states of the opposite branch at the same backbone layer. The resulting guidance modulates the sequence-level advantage to obtain token-level advantages, allowing more precise policy updates. Experiments on Janus-Pro and Show-o, covering pure autoregressive and hybrid autoregressive–diffusion architectures, show consistent improvements over joint GRPO in multimodal understanding, image captioning, and text-to-image generation. Further gradient analysis reveals increased cross-branch gradient agreement. These results demonstrate that CGRL reduces the optimization conflict between understanding and generation and enables mutually beneficial joint RL in unified multimodal models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.