SimHack: Predicting the onset of detectable reward hacking before RL training
Abstract
While reinforcement learning can substantially improve language models at coding, it can induce reward hacking, where models exploit imperfections in the reward signal rather than solving the intended task. Once reward hacking is detected, existing work largely focuses on environment hardening, which becomes increasingly difficult as training environments grow more complex and models become more capable. We instead study whether the onset of detectable reward hacking can be predicted before executing a full RL run, thus allowing us to proactively avoid reward hacking. We formulate this problem as predicting the distribution of reward hacking onset steps under a target training recipe. We introduce SimHack, a lightweight behavioral simulator that models RL as a stochastic behavioral process at the prompt level. SimHack learns these behavior dynamics from a handful of smaller-scale reference RL runs and generates counterfactual simulations for unseen target recipes without costly retraining. To support further research on reward hacking prediction, we publish a dataset of 325 coding RL runs that spans diverse training recipes and reward hacking behaviors. Across unseen target recipes, SimHack achieves a median onset relative error of 11.3% and an across-recipe rank correlation of 0.85, while transferring across changes in model scale, curriculum, number of rollouts, and learning rate. Using these predictions, we show that SimHack can guide training decisions by screening candidate curricula and identifying prompts that disproportionately contribute to the propagation of reward hacking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.