Rubric Can Hallucinate Too: Diagnosing and Repairing Multimodal Rubric Hallucination
Abstract
As multimodal large language models (MLLMs) are deployed in increasingly open-ended settings, scalar correctness and holistic preference scores provide limited diagnostic feedback. Evaluation and alignment are therefore shifting toward rubric-based assessment, which decomposes response quality into explicit, scoreable criteria. However, these criteria are not always reliable: they may demand information absent from the image, assert unverifiable motives, or require deliverables beyond what the task asks for. We term this phenomenon **Multimodal Rubric Hallucination (MRH)** and formalize rubric validity along two axes: Evidence Support and Task Fit. Unlike conventional multimodal hallucination in model responses, MRH makes the evaluation standard itself the object of diagnosis. We propose a new task, **MRH Diagnosis and Repair**: identify one of six MRH types and rewrite the hallucinated rubric into an evidence-supported, task-aligned replacement. We introduce **MRH-Bench**, comprising 8,472 verified multimodal instances organized into three progressively challenging tracks: rubric hallucination type identification, hallucinated rubric repair, and LLM–human evaluation agreement assessment. Across eight representative proprietary and open-weight MLLMs, the strongest prompted baseline reaches only 61.0% Macro-F1 and 15.80% repair success. We further propose **MRH-Guard**, a two-stage framework combining structured weighted supervised fine-tuning with **Quadrant-Aware Policy Optimization (QAPO)** to strengthen the diagnostic decisions that guide repair. MRH-Guard reaches 75.4% Macro-F1 (+14.4 points) and 19.20% repair success (+3.40 points).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.