acceptodds
Under review as a conference paper at ICLR 2027

LOCAL: Local-Curvature Advantage Logit Regression for Scalable Policy Optimization

Abstract

Policy optimization should account for how parameter updates change the policy distribution rather than their magnitude in parameter space. Natural policy gradient (NPG) and policy mirror descent (PMD) achieve this through parameter-space and state-wise curvature, respectively, and yield the same Boltzmann policy update for tabular softmax policies. This update determines the target logits only up to a state-dependent offset (SDO), since adding a common constant to all logits leaves the policy unchanged. While therefore negligible in the tabular setting, we show that the SDO becomes consequential under function approximation: fixing it during target fitting alters the curvature beyond the local Fisher geometry. We introduce LOCAL (**Lo**cal-**C**urvature **A**dvantage **L**ogit Regression), which analytically profiles out the SDO and regresses only the policy-relevant relative logit changes. We show that the resulting objective preserves the local Fisher curvature in logit space and that its local linearization recovers NPG. We further derive a first-order implementation and a Top- approximation that makes the method practical for very large action spaces. Across Atari and Procgen controls and LLM post-training with RLHF and RLVR, LOCAL consistently improves policy optimization, showing that properly handling this seemingly redundant degree of freedom is important for policy learning with function approximation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.