Don’t Overthink: Adaptive Thinking that Does not Fireback for Image Editing using GRPO
Abstract
Reasoning-augmented image editing operates under a prevailing assumption: more deliberation yields better edits. We show this assumption is fundamentally flawed. Through a diagnostic study of BAGEL, a state-of-the-art unified multimodal model, we identify three root pathologies of SFT-based reasoning: complexity-agnostic outputs that apply identical deliberation depth to trivial and complex edits alike; adherence dependency, where correct reasoning fails to translate into generation gains without sufficient adherence capacity in the generation expert; and verbosity without substance, where tripling the reasoning budget yields no quality improvement (GPT-Score ≈5.2), while structural directives at the same budget deliver a substantial gain to 6.1. These findings converge on a single conclusion: neither adaptive reasoning length nor structured templates alone resolves overthinking, both are jointly necessary. We introduce PRISM (Planning with Reasoning via Instructed Structured Modes), a compound adaptive-and-structured reasoning framework that decomposes pre-generation deliberation into three sequential stages, complexity steering, tailored lens generation, and ordered lens execution, distilled into unambiguous per-sample actions that drive a flow-matching generation expert. The understanding and generation experts are trained jointly via GRPO and Flow- GRPO with variance-guided early-window selection, without any external verifier or multi-turn inference. To power the framework, we introduce REVISE-50K, a dataset of 50,000 samples spanning 50 editing skills with graded complexity labels and disentangled hardness and verbosity axes. PRISM establishes a new state of the art on three reasoning-intensive benchmarks, +13.0% on KRIS, +3.4% on RISE, +17.8% on PICA, while preserving performance on direct editing tasks and surpassing multi-round external verifier systems at 20× lower inference cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.