acceptodds
Under review as a conference paper at ICLR 2027

Unified Reinforcement Learning with Progressive Visual Reasoning for Text-Aware Image Restoration

Abstract

Text-Aware Image Restoration is inherently challenging because it is a severely ill-posed problem where character strokes are extremely sensitive to distortion. While recent methods integrate Vision-Language Models for guidance, they often suffer from a VLM drift phenomenon where the guidance policy sacrifices OCR accuracy to satisfy the generative biases of a frozen restoration backbone. In this paper, we propose a novel unified reinforcement learning framework that enables the co-evolution of a VLM-based guidance policy and a Diffusion Transformer restoration model. We formulate the restoration process as an active exploration within the rich pre-trained search space of the generative model to find optimal, text-faithful renderings among multiple plausible solutions. To resolve the non-stationary transition dynamics of this joint environment, our framework employs an Asymmetric Hybrid Reinforcement Learning strategy. Specifically, the VLM is optimized via an on-policy approach using Group Relative Policy Optimization, while the Diffusion Transformer is updated via an off-policy approach using DiffusionNFT. Furthermore, we introduce a Progressive Search mechanism that utilizes intermediate denoised estimates and real-time feedback along the denoising trajectory to iteratively refine the textual search space. Extensive experiments on the SA-Text and Real-Text benchmarks demonstrate that our framework achieves new state-of-the-art performance by effectively reducing hallucinations and significantly enhancing functional legibility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.