acceptodds
Under review as a conference paper at ICLR 2027

CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL

Abstract

During Reinforcement Learning with Verifiable Rewards (RLVR), large language models may exploit loopholes in their environments to earn rewards instead of completing the intended tasks, a phenomenon known as reward hacking. We introduce LoopholeRL, a small, controllable testbed for studying reward hacking in coding RL. We first turn algorithmic problems into software engineering tasks with deliberately exposed loopholes, then control hacking through two key elements of RL: the base model and reward design. We use synthetic data to perform Supervised Fine-Tuning (SFT) on Qwen3-4B, producing base models with different hacking tendencies, and shape rewards by adjusting the relative weights of weak and strong tests to vary the difficulty of earning rewards during RL. The framework also supports evaluating reward hacking mitigation methods. Our experiments show that higher proportions of hacking demonstrations in SFT data and greater reward weight on strong tests advance hacking onset. Among the tested mitigation methods, a Chain-of-Thought (CoT) monitor achieves the highest coding performance at the evaluated checkpoint, but its initial suppression of hacking does not persist. Notably, under monitoring pressure, the model's code comments increasingly mislead the monitor: removing them makes reward hacking easier to detect.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.