Cross-Case Rubric Learning for Consistent Visual Preference Evaluation
Abstract
Multimodal large language models (MLLMs) can produce plausible rubric-based assessments of visual creatives, yet these often lack cross-case consistency and fail to translate into accurate preference evaluations. To bridge this gap, we introduce a curated dataset of 7,044 same-product image pairs, with preference labels derived directly from aggregate click-through rates (CTR) in real user exposure logs. We then propose Cross-Case Rubric Learning (CCRL), a framework that connects rubric-based visual reasoning with these behavioral outcomes. CCRL decouples preference evaluation into identifying visual differences, selecting relevant criteria, and assigning dimension-level scores that are deterministically aggregated. We optimize this scoring with a leave-one-out policy-gradient (RLOO) estimator, combining local rewards with conditional cross-case feedback. Specifically, independently assessed visual relationships act as relational constraints, penalizing inconsistent scores for equivalent visual elements across different cases. CCRL achieves the highest reported pairwise accuracy among the evaluated baselines. We assess intermediate evaluation quality using expert-reference rubric micro-F1 and full-program compliance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.