acceptodds
Under review as a conference paper at ICLR 2027

AxisRubric: Rubric Construction via Axis Discovery for Judging Multimodal Generation

Abstract

Rubric-based rewards improve comparative preference prediction for multimodal generation by decomposing human preferences into explicitly checkable evaluation axes. Their accuracy therefore depends on how effectively these axes capture the factors that shape human preference. A largely overlooked challenge in discovering preference-relevant axes for reward modeling in multimodal generation is the cross-modal observability gap. That is, textual instructions and predefined rubrics may miss delivery-quality differences, such as rendering artifacts, audio clarity, and motion coherence, which are observable only in the generated multimodal outputs. Existing methods use instruction-derived, expert-authored, or output-aware rubrics, yet none fully closes the gap. Our preliminary study shows that even when the rubric author sees the generated outputs, it still focuses on what the instruction requests, and up to 32% of human-validated delivery-quality axes are not discovered. Moreover, the discovered axes are not equally predictive of human preference, nor are they applicable to every instance. These observations reveal two key failure modes: axis omission, where preference-relevant axes are missing from the rubric; dilution, where redundant, poorly predictive, or inapplicable axes weaken the influence of informative ones. To close the gap, we propose AxisRubric, a three-phase multi-agent framework for automatic rubric construction via axis discovery. To address omission, Phase 1 runs independent instruction-blind and instruction-informed observer agents that discover evaluation axes, which we consolidate into an evolving benchmark-level taxonomy. To address dilution, Phase 2 deactivates misleading axis categories and calibrates the weights of the remaining axes, and Phase 3 selects the axes applicable to each individual instance and rewrites them as instance-specific rubric items. Across twelve preference benchmarks spanning image, audio, and video generation, AxisRubric outperforms all seven evaluated baselines. It improves preference accuracy over the strongest automatic rubric construction baseline by 0.08 on average (from 0.66 to 0.74) and by up to 0.18 (from 0.70 to 0.88). Notably, these gains hold across 21 pairs of rubric-author and output-judge models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.