Heat-GRPO: Heat-Kernel Group Relative Policy Optimization for Multi-Rubric Reinforcement Learning
Abstract
Training language models to generate useful and reliable responses requires feedback that captures multiple dimensions of quality beyond correctness. Multi-rubric reinforcement learning offers such feedback, yet reducing nuanced judgments to discrete rewards can discard the distinctions needed for learning. We empirically analyze how nonlinear scoring of rubric-level judge margins before aggregation changes group-relative learning. Heat-GRPO realizes this design choice by smoothing hard verdicts under Gaussian margin perturbations, preserving margin distinctions while limiting the influence of extreme judgments. Identical-response comparisons isolate ranking changes; controlled training comparisons under shared scalar normalization demonstrate gains over Hard and Linear scoring. Comparing complete learning configurations across six models—Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Llama-3.2-1B, Llama-3.2-3B, and Llama-3.1-8B—and four benchmarks, Heat-GRPO improves rubric scores by up to 25.0% relative to the strongest baseline and achieves the highest three-seed mean in 23 of 24 conditions. Vectorized scoring and scalar aggregation reduce advantage computation time by 38.9% and 38.2% relative to GDPO and GD2PO, respectively, while retaining a cost comparable to GRPO. Additional evaluation with Qwen3.8-27B, gpt-oss-20b, and Gemma-4-31B ranks Heat-GRPO first in 22, 20, and 21 of the 24 conditions, respectively, supporting the robustness of its overall advantage across evaluators.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.