Prior Reward Hacking Impairs Subsequent Learning in Exploitable Environments
Abstract
In reinforcement learning, reward hacking occurs when a large language model exploits weaknesses in an evaluator to obtain reward without completing the intended task. Although reward hacking has been shown to transfer to broader misalignment at evaluation, its effects on a model's subsequent learning trajectories remain poorly characterized. We show that reward hacking behavior acquired during training can impair later learning in new domains: in our experiments, we first expose models to small amounts of misaligned training data that reward exploiting the training objective rather than solving the task as intended. During subsequent post-training in new domains, these models systematically favor discovery and exploitation of shortcuts, failing to acquire the intended task-solving behavior, while models trained without misaligned rewards reliably learn to solve the tasks as intended. In our primary experimental setting, this occurs even when exploiting the shortcut yields less reward than solving the task correctly. We observe this phenomenon across multiple downstream domains, model families, and task types. In a more limited experiment, we further show this impairment can extend to a later environment with no exploit available at all, once shortcut convergence has already taken hold. We additionally release hundreds of reinforcement-learning tasks spanning coding, finance, and life sciences to support reproducibility and further study. Our work uncovers a vulnerability in multi-stage LLM post-training: reward hacking learned early in training can impair the acquisition of intended behaviors in later, distinct training environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.