RubricVerifier-STEM: Auditable Criterion-Level Rewards for STEM Reasoning
Abstract
Rubric-based rewards offer a promising way to provide dense supervision for longform STEM reasoning, but their effectiveness ultimately depends on the verifier that converts rubric criteria into reward. We study rubric verifiers as criterionvector reward interfaces and evaluate three properties required for their use in reinforcement learning: criterion-verdict correctness, surface-form invariance, and servable output compliance. To study these requirements, we introduce RUBRICHARD-STEM, a 500-example diagnostic benchmark spanning five STEM domains and 4,149 criterion labels, and develop RUBRICVERIFIER-STEM, a compact 4B verifier trained with diverse-source SFT, agreement-first RSFT, and interfaceaware RL. Despite its scale, RUBRICVERIFIER-STEM reaches 84.79% criterion accuracy and 83.94% per-example macro accuracy, outperforming all evaluated specialized verifier baselines, while using only 27% as many completion tokens as GPT-OSS-120B. Using RUBRICVERIFIER-STEM as the reward model raises Qwen3-4B-Thinking’s average score across 17 STEM benchmarks from 38.26% to 47.52%. The resulting score is 0.42 percentage points below the GPT-OSS-120B reward run. Together, these results show that a compact, auditable criterion-level verifier can provide an effective reward interface for STEM reasoning, connecting local grading fidelity to downstream RL performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.