acceptodds
Under review as a conference paper at ICLR 2027

RubricEvo: Escaping Uniform Rewards in Rubric-based Text-to-Image Reinforcement Learning via Online Rubric Evolution

Abstract

Rubric-based rewards have recently gained attention in text-to-image RL, as decomposing prompt-following into atomic, verifiable criteria enables comprehensive verification of complex compositional semantics. However, existing rubric-based methods for text-to-image RL construct their rubric sets before training and keep them fixed throughout training. The fixed rubrics are never checked against the images that the current policy generates, so their difficulty can mismatch the ability of the policy. Specifically, many rubrics are uniform: they are satisfied or failed by nearly all rollouts. A uniform rubric assigns nearly the same score to every image in a rollout group. Rewards then carry little difference across samples, leaving the policy with almost no guidance on which samples to reinforce. To address this, we propose , an online rubric evolution framework that revises rubrics at each training step based on feedback from the policy's current rollouts, keeping the rubrics discriminative throughout training. The feedback takes two forms. The pass rate of each rubric identifies the uniform ones and tells the direction to adjust their difficulty. The visual differences among rollout images expose distinctions that the existing rubrics do not cover, from which new rubrics are constructed. Since the two forms of feedback are heterogeneous, handling them in a single MLLM call risks letting the rubric text dominate the visual evidence. RubricEvo therefore processes them with a two-stream strategy, yielding rubrics both faithful to the prompt and grounded in the actual differences among rollouts. The rollouts are then re-scored under the updated rubrics to restore discriminative rewards. Extensive experiments on widely adopted text-to-image benchmarks show that RubricEvo consistently outperforms state-of-the-art scalar-based and rubric-based methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.