acceptodds
Under review as a conference paper at ICLR 2027

RHDP: A Gradient-Variation-Aware Reward Hacking Detection and Penalization Mechanism for Chest X-ray Report Generation

Abstract

Chest X-ray report generation aims to produce a detailed and clinically accurate description for a given image. Although reinforcement learning (RL) fine-tuning can optimize sequence-level preferences, an imperfect proxy reward may be over-optimized, leading to template collapse and deterioration of the clinically meaningful gold objective. We introduce RHDP, a reward-hacking detection and penalization framework built around the long-window variance of input gradients (LVoG). LVoG is the temporal variance of the absolute input gradient of the joint CE–RL loss and is used as a sample-wise reliability score: abnormally low values indicate increased proxy–gold misalignment risk. A decomposition analysis shows that LVoG combines CE–RL transition dynamics, reward-side advantage dynamics, and pure RL policy-side image-conditioned sensitivity; the policy-side signal remains predictive after removing instantaneous advantage scaling. RHDP uses lagged LVoG to modulate the reward and the forward KL divergence, suppressing unreliable updates while preserving generation quality. Experiments on IU-Xray and MIMIC-CXR demonstrate improved clinical accuracy, competitive semantic quality, and consistent mitigation of reward hacking. The code is available on GitHub.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.