acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Self-Correction for Multi-Turn Interleaved Text-Image Generation

Abstract

Self-correction in interleaved text-image generation requires deciding when to refine an image, fall back to a new attempt, or stop, based on the evolving generation history. We introduce *Adaptive Self-Correction Policy Optimization* (ASPO), an on-policy reinforcement learning framework that learns a variable-length policy over reasoning, correction actions, and image synthesis. Cold-start supervision establishes the correction behaviors; reinforcement learning then jointly optimizes the language-action and image-generation branches using only final-image rewards. Each intermediate image provides context for subsequent decisions. For the rectified-flow image branch, ASPO replays one cached stochastic denoising transition per generated image, yielding a tractable transition-level policy surrogate for optimization from the final trajectory outcome. We also introduce *CorrectionBench*, comprising 350 instances that evaluate recovery from flawed interleaved contexts. ASPO improves overall generation performance on GENIUS and GenEval++ over supervised initialization. On CorrectionBench, it achieves gains of 7.9 points in automated evaluation and 7.53 points in blinded human evaluation over the same baseline. Behavioral analyses and component ablations further support the learned correction behavior and the contribution of image-policy optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.