acceptodds
Under review as a conference paper at ICLR 2027

GAUGE: Gauge-Invariant Policy Optimization for Language Agents

Abstract

Policy optimization for language agents typically constrains changes in token likelihood, although environment transitions and rewards depend on executable actions obtained by parsing the generated text. This mismatch can over-penalize different textual renderings of the same action while under-resolving small but behaviorally critical changes in action arguments. We formalize this distinction by decomposing raw-text policy drift into drift over executable actions and within-action gauge drift. Based on this decomposition, we introduce GAUGE, a parser-aware trust-control method that aggregates likelihood across candidate renderings and measures drift at both the action and semantic-slot levels. GAUGE preserves sequence-level gradient estimation and is compatible with group-relative policy optimization. We evaluate it in the ARLArena pipeline across ALFWorld, WebShop, Sokoban, and TIR Math. With SFT-initialized Qwen3-4B, GAUGE achieves 75.09% success on WebShop and 85.67% on Sokoban, improving over the corresponding GSPO and DAPO baselines by 2.61 and 3.27 percentage points. These results support executable-action space as a useful representation for controlling language-agent policy updates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.