acceptodds
Under review as a conference paper at ICLR 2027

Learning to Repair: Skill Optimization for Multi-Round Text-to-Image Generation

Abstract

Multi-round text-to-image generation uses feedback to correct compositional errors, yet additional refinement does not always lead to better images. Progress depends on choosing suitable repair actions and ordering them so that later operations preserve earlier corrections. How can repair experience be turned into reusable guidance for these decisions? We address this question through skill optimization, with a focus on designing the space of guidance to be learned. Different failure types respond differently to editing and regeneration, while interactions between repairs make their order consequential. Motivated by these observations, we structure the critic's natural-language skill around failure diagnosis, type-conditional action choice, preservation requirements, and repair priorities, with generator-specific observations stored separately. Starting from human-designed guidance, we use feedback from a small collection of historical trajectories to refine the skill while keeping the underlying models frozen. Experiments on GenEval, GenEvalĀ 2, and ConceptMix show improvements over the IRG baseline. Removing either action or ordering guidance reduces the effectiveness of the optimized skill, while transfer from GenEval2 to ConceptMix shows that the learned guidance retains value beyond its source dataset. These findings support structured skill optimization as a way to improve compositional generation by learning how to repair.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.