Cycle Alignment for Understanding-Guided Visual Editing in Unified Multimodal Models
Abstract
Image editing demands both fine-grained visual understanding and controllable generation. A model must comprehend objects and layouts in the source image, parse editing instructions, and modify specific regions while preserving identity and background. Unified multimodal models (UMMs) inherently possess both capabilities within a single architecture, making them well-suited for editing. However, in conventional UMM training, the understanding and generation branches are optimized independently. Despite sharing the same backbone, the potential for understanding to directly guide generation remains largely unexploited. To exploit this potential, we propose Cycle Alignment (CycleA), a post-training paradigm that closes the loop between understanding and generation. Specifically, CycleA routes the generation branch's output back into the frozen understanding branch, where source-image VQA constraints provide differentiable semantic reconstruction feedback to the generation branch. This feedback operates directly in its native feature space through a lightweight frozen projector, avoiding the decode-encode overhead. To prevent the mapped features from drifting, we further introduce a feature alignment loss that keeps them close to real visual features. Across the evaluated configurations, CycleA improves aggregate image-editing scores over BAGEL after a 2,000-step post-training schedule following 5,000-step projector pretraining, while text-to-image benchmark scores remain comparable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.