acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Policy Trust Region: Attributing and Controlling Representation-Induced Movement in PPO

Abstract

Trust-region methods constrain how far a policy moves, but they do not reveal which part of the policy network produces that movement. In policies with a learned representation and an action head, updates to the two components can reinforce or compensate for one another, so the complete old-to-new policy change can hide substantial component-level movement. We make these dynamics observable by recombining old and new representations with old and new action heads to form four counterfactual policies. In particular, evaluating the new representation through the frozen behavior head measures the action-distribution change attributable to the representation update. Building on this construction, we introduce the Movement-Attributed Representation Trust Region (\MART), a plug-in regularizer that limits excess movement along this representation path while preserving the underlying PPO-family actor objective. Across extensive experiments in open source environments from various control scenarios, \MART reduces the deviation of the attributed movement to the representation and the action ratio of the sample in all the settings tested. In traffic control, these reductions become stronger as each rollout is reused for more optimization epochs, and the action head often moves farther when the representation path is constrained, revealing compensation that is partially hidden by complete-policy diagnostics. Moreover, the results show that controlling policy movement generally benefits from considering not only the complete policy update, but also the component through which that movement arises.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.