APD-Flow: Reward Post-Training for Flow-Matching Image Restoration
Abstract
Flow-matching image restorers produce high-quality outputs with deterministic inference, but they are trained to match paired ground truth rather than human preferences or task-specific rewards. Reward post-training offers a way to close this gap. However, policy-gradient methods rely on stochastic exploration and action likelihoods, neither of which a deterministic sampler provides. We propose *APD-Flow*, a reward post-training framework that introduces exploration at the source state instead of the sampling trajectory. Given a degraded input, *APD-Flow* perturbs the source state in two opposite directions and restores both with the same deterministic sampler. The difference between the two restorations, weighted by their reward difference, defines an Antithetic Preference Direction (APD) that points toward the preferred output. The restorer is then optimized along this direction without action likelihoods or trajectory noise, and fidelity-aware reward adaptation keeps fidelity at the level of the pretrained restorer. The framework applies to restorers with arbitrary numbers of sampling steps and to non-differentiable rewards. Experiments on blind face restoration, blind image super-resolution, and scene-text restoration show that *APD-Flow* consistently improves the targeted perceptual or task-specific metrics while keeping the inference architecture and budget unchanged. Code and a video explanation are provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.