acceptodds
Under review as a conference paper at ICLR 2027

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

Abstract

Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images according to prompt alignment and perceptual quality. Existing reward models commonly require training on large-scale human preference corpora. Meanwhile, more recent work leverages Vision-Language Model (VLM) judges with rubrics to make evaluation criteria explicit. However, rubric-based approaches often lack reliable methods for criterion selection. In this paper, we propose , a framework that automatically synthesizes, selects, and weights explicit rubrics to guide VLM judges. AutoRubric-T2I first extracts candidate rubrics from reasoning traces over preference pairs and scores the paired images under each rubric using a frozen VLM judge. Then, an \textbf{\ell_1-Regularized Logistic Regression Refiner} fits these scores to human preference labels, pruning rubrics that fail to guide the judge toward preference-consistent assessments and learning dataset-specific weights for the retained criteria. Using just 256 preference pairs (0.02-0.04% of full RM training data) during training, AutoRubric-T2I outperforms fine-tuned reward models on MMRB2. As an RL reward in Flow-GRPO, it also improves generation quality over the fine-tuned reward models on TIIF and UniGenBench++.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.