CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
Abstract
During Reinforcement Learning with Verifiable Rewards (RLVR), large language models may exploit loopholes in their environments to earn rewards instead of completing the intended tasks, a phenomenon known as reward hacking. We introduce LoopholeRL, a small, controllable testbed for studying reward hacking in coding RL. We first turn algorithmic problems into software engineering tasks with deliberately exposed loopholes, then control hacking through two key elements of RL: the base model and reward design. We use synthetic data to perform Supervised Fine-Tuning (SFT) on Qwen3-4B, producing base models with different hacking tendencies, and shape rewards by adjusting the relative weights of weak and strong tests to vary the difficulty of earning rewards during RL. The framework also supports evaluating reward hacking mitigation methods. Our experiments show that higher proportions of hacking demonstrations in SFT data and greater reward weight on strong tests advance hacking onset. Among the tested mitigation methods, a Chain-of-Thought (CoT) monitor achieves the highest coding performance at the evaluated checkpoint, but its initial suppression of hacking does not persist. Notably, under monitoring pressure, the model's code comments increasingly mislead the monitor: removing them makes reward hacking easier to detect.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.