acceptodds
Under review as a conference paper at ICLR 2027

SIGHT: Compositional Skill Routing for Instance-Specific Rubric Generation across Heterogeneous Multimodal Tasks

Abstract

Reward models guide reinforcement learning (RL) for multimodal large language models with scalar scores that obscure multi-dimensional human judgment criteria. To address this, rubric-as-reward decomposes response quality into structured, verifiable criteria. Despite rapid adoption, two questions remain: can benchmarks assess rubric fidelity to human-valued dimensions, and how can prompt-specific rubrics reflect human preferences across heterogeneous multimodal tasks? We introduce **MM-RubricBench**, the first multimodal rubric-guided reward benchmark, spanning six heterogeneous domains. Each curated question has a weighted expert-authored rubric and human ratings of eleven models' responses scored against it. These annotations support evaluation of judge-human preference alignment and generated rubric quality. We further propose **SIGHT**, compositional **S**kill routing for **I**nstance-specific rubric **G**eneration across **H**eterogeneous multimodal **T**asks. SIGHT transfers construction logic rather than rubric content: it extracts reusable skills from expert rubrics, builds instance-specific evaluation blueprints, selects complementary skills to cover them, and generates rubrics traceable to specific requirements. On MM-RubricBench, these rubrics outperform model-based scoring protocols in human-preference alignment, approaching the human rubric ceiling. When used as RL rewards, these rubrics also improve policies over baselines in multimodal reasoning, instruction following, and agentic rewarding. Construction-skill routing thus supports practical, scalable, and reliable multimodal alignment. Code, datasets, and benchmark will be publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.