DiT-Reward: Generative Representations for Text-to-Image Reward Modeling
Abstract
Can representations learned for image generation also support the evaluation of generated images? We study text-to-image reward prediction as a downstream task of generative representation learning and introduce DiT-Reward, which repurposes pretrained diffusion transformers by aggregating image representations conditioned on text across layers. Probing a frozen SD3.5-Large backbone reveals that preference information is already accessible through a lightweight learned head, with stronger predictive features in the middle and later layers and complementary information across stages. Using the same public preference data mixture as HPSv3, both our SD3.5-Large and Qwen-Image variants outperform HPSv3 on all four evaluated benchmarks, with the Qwen-Image variant reaching 79.1% accuracy on HPDv3. Backbone comparisons further indicate a positive scaling trend with generative backbone capacity. When used as the reward for Flow-GRPO, the SD3.5-Large variant produces a policy that wins 75.1% of 2,048 human comparisons against the policy optimized with HPSv3, with 21.6% losses and 3.3% ties, after aggregating five ratings per image pair. With a shared latent space between the policy and reward model, direct latent scoring also achieves a 1.65× reward inference speedup over HPSv3 with comparable peak memory. These results establish generative representations as an effective foundation for reward modeling and policy optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.