RHContrast: A Matched-Control Benchmark for Evaluating Reward-Hacking Risks in Long-Form LLM Responses
Abstract
Reward changes after rewriting can be misleading: an erroneous answer may gain reward while benefiting less than a reference answer, or lose reward while incurring a smaller penalty than the reference. We introduce FAME-RHBench, a long-form benchmark that tracks reference answers and variants with localized defects through matched transformations, preserving explicit links between original and transformed texts. The benchmark separates reward recovery, changes in the reference–defect score gap, and direct comparison with an untransformed reference. Naturalization reorganizes both answer branches; verification endorsements add the same unsupported claim without altering either answer body. Across seven prompted judge families, naturalization increases scores on both branches on average, but benefits reference answers more. Verification endorsements lower both scores, yet penalize reference answers more strongly. Analyses with scalar and pairwise reward models, grounding verifiers, and an AI-assisted, human-reviewed reference characterize how these responses vary across evaluation settings. The benchmark enables evaluator audits that separate general transformation effects from defect-conditioned score changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.