acceptodds
Under review as a conference paper at ICLR 2027

DiT-Reward: Generative Representations for Text-to-Image Reward Modeling

Abstract

Can representations learned for image generation also support the evaluation of generated images? We study text-to-image reward prediction as a downstream task of generative representation learning and introduce DiT-Reward, which repurposes pretrained diffusion transformers by aggregating image representations conditioned on text across layers. Probing a frozen SD3.5-Large backbone reveals that preference information is already accessible through a lightweight learned head, with stronger predictive features in the middle and later layers and complementary information across stages. Using the same public preference data mixture as HPSv3, both our SD3.5-Large and Qwen-Image variants outperform HPSv3 on all four evaluated benchmarks, with the Qwen-Image variant reaching 79.1% accuracy on HPDv3. Backbone comparisons further indicate a positive scaling trend with generative backbone capacity. When used as the reward for Flow-GRPO, the SD3.5-Large variant produces a policy that wins 75.1% of 2,048 human comparisons against the policy optimized with HPSv3, with 21.6% losses and 3.3% ties, after aggregating five ratings per image pair. With a shared latent space between the policy and reward model, direct latent scoring also achieves a 1.65× reward inference speedup over HPSv3 with comparable peak memory. These results establish generative representations as an effective foundation for reward modeling and policy optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.