acceptodds
Under review as a conference paper at ICLR 2027

RHContrast: A Matched-Control Benchmark for Evaluating Reward-Hacking Risks in Long-Form LLM Responses

Abstract

Reward changes after rewriting can be misleading: an erroneous answer may gain reward while benefiting less than a reference answer, or lose reward while incurring a smaller penalty than the reference. We introduce FAME-RHBench, a long-form benchmark that tracks reference answers and variants with localized defects through matched transformations, preserving explicit links between original and transformed texts. The benchmark separates reward recovery, changes in the reference–defect score gap, and direct comparison with an untransformed reference. Naturalization reorganizes both answer branches; verification endorsements add the same unsupported claim without altering either answer body. Across seven prompted judge families, naturalization increases scores on both branches on average, but benefits reference answers more. Verification endorsements lower both scores, yet penalize reference answers more strongly. Analyses with scalar and pairwise reward models, grounding verifiers, and an AI-assisted, human-reviewed reference characterize how these responses vary across evaluation settings. The benchmark enables evaluator audits that separate general transformation effects from defect-conditioned score changes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.