acceptodds
Under review as a conference paper at ICLR 2027

Reward Moves What It Measures: Anisotropic Program Anchoring for RL Fine-Tuning under Incomplete Objectives

Abstract

Reinforcement learning fine-tuning optimizes generative policies against objectives that are incomplete: a reward scores what it can measure, never everything the designer wants, and optimization then degrades the unmeasured behavior. The standard remedy is a scalar KL penalty toward a frozen clone of the initial policy. This remedy has two limitations: it holds every behavioral dimension equally, so it can only slow the whole policy, and its reference is a snapshot that loses validity once the policy leaves the states it was fitted on. We propose Anisotropic Program Anchoring (APA), which anchors the policy to an executable symbolic program instead. The program encodes the unmeasured behavior directly and can be queried at any state the policy reaches, and one anchor coefficient per behavioral dimension lets the reward move what it measures while the program holds the rest. We evaluate APA on a musical task where the objective is incomplete by nature: a piano and a bass agent fine-tuned together to accompany a jazz soloist. In a factorial design, we compare ten experimental conditions on held-out measures of harmonic competence that the reward never scores. Without anchoring, competence falls by half. APA keeps competence within 1% of its initial value while realizing 80% of the unanchored reward gain, against about 55% for scalar anchors. Under distribution shift it loses 4% of competence where an anchor to a frozen clone loses 38%. Adding another reward term for the most visible failure repairs only that property, while others keep degrading.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.