: Repairing Only What Misleads the Policy from Sparse Outcomes
Abstract
Dense rewards guide learning well but are often misspecified, leading to reward hacking. Sparse trajectory outcomes are reliable but provide weak guidance. Existing methods either learn a new reward from feedback or restrict how far the policy departs from reference behavior. We take a different route and repair the dense proxy itself, using sparse outcomes to correct the errors that cause hacking while keeping its guidance. The difficulty is that an outcome only measures the total reward of a trajectory, so a correction can fit every outcome and still favor the wrong policy. We quantify this ambiguity with an ambiguity modulus and use it to bound the regret of the repaired policy. The analysis shows that the ambiguity arises in two ways: errors can cancel out within a trajectory, or lie on behavior the data never covers. Fitting the outcomes more accurately fixes neither. We therefore propose Adaptive Identification for Reward Repair (), which queries outcomes on trajectory prefixes to catch errors that cancel, and collects outcomes from an exploration policy to cover what the data missed. also estimates a potential function alongside the correction, so that harmless shaping in the proxy is not mistaken for error. The final correction is frozen and passed to a standard RL learner. On 22 continuous control tasks with constructed reward misspecification, instantiated with BAC achieves the highest aggregate normalized recovery score among the evaluated methods. We further apply with PPO to three published benchmarks and observe substantial improvements over the baseline methods in both experimental settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.