Reward Moves What It Measures: Anisotropic Program Anchoring for RL Fine-Tuning under Incomplete Objectives
Abstract
Reinforcement learning fine-tuning optimizes generative policies against objectives that are incomplete: a reward scores what it can measure, never everything the designer wants, and optimization then degrades the unmeasured behavior. The standard remedy is a scalar KL penalty toward a frozen clone of the initial policy. This remedy has two limitations: it holds every behavioral dimension equally, so it can only slow the whole policy, and its reference is a snapshot that loses validity once the policy leaves the states it was fitted on. We propose Anisotropic Program Anchoring (APA), which anchors the policy to an executable symbolic program instead. The program encodes the unmeasured behavior directly and can be queried at any state the policy reaches, and one anchor coefficient per behavioral dimension lets the reward move what it measures while the program holds the rest. We evaluate APA on a musical task where the objective is incomplete by nature: a piano and a bass agent fine-tuned together to accompany a jazz soloist. In a factorial design, we compare ten experimental conditions on held-out measures of harmonic competence that the reward never scores. Without anchoring, competence falls by half. APA keeps competence within 1% of its initial value while realizing 80% of the unanchored reward gain, against about 55% for scalar anchors. Under distribution shift it loses 4% of competence where an anchor to a frozen clone loses 38%. Adding another reward term for the most visible failure repairs only that property, while others keep degrading.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.