VideoRubric: Rubric Granularity Matters in Reinforcement Learning for Video Reasoning
Abstract
Reinforcement learning with verifiable rewards is promising for improving video reasoning in multimodal large language models, yet answer-level rewards provide limited guidance on the reasoning process. We introduce VideoRubric, a rubric-guided reinforcement learning framework that studies process supervision for VideoQA at two granularities: specific rubrics describe sample-conditioned evidence and reasoning requirements, while abstract rubrics capture reusable patterns such as temporal consistency, causal modeling, and evidence grounding. VideoRubric constructs specific rubrics with a multimodal teacher, abstracts them into a shared rubric set, and incorporates rubric judgments into GRPO through correctness gating. Experiments with Qwen2.5-VL-7B-Instruct on 1K rubric-augmented samples show that the correctness-gated specific-rubric supervision achieves the strongest average performance across MMVU, TempCompass, VSI-Bench, and VideoMME. Further analyses show that rubric granularity shapes learned reasoning differently: specific rubrics yield clearer separation between correct and incorrect responses, whereas abstract rubrics produce more structured and preferred reasoning outputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.