ReasonUni: Reasoning-Augmented Reconstruction Advances Unified Multimodal Models
Abstract
Image-to-text (I2T) understanding and text-to-image (T2I) generation are dual tasks in unified multimodal models (UMMs). Reconstructive post-training couples these tasks, but existing reconstruction-based methods remain perception-centric and provide limited supervision for reasoning about structure, intent, and transformation. We propose ReasonUni, a reasoning-augmented reconstructive post-training framework that extends this coupling beyond perception. Given an original image and an editing instruction, the UMM first reasons about what and how to edit and describes the target state, then generates the corresponding target image conditioned on this understanding output. To jointly optimize both modules, we introduce TreeGRPO, a tree-structured reinforcement learning method that samples multiple reasoning paths and candidate images per path, then aggregates visual rewards to update both modules under a unified objective. Experiments across multiple benchmarks demonstrate consistent gains in both understanding and generation, showing that reasoning-augmented reconstruction yields genuine mutual gains for UMMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.