acceptodds
Under review as a conference paper at ICLR 2027

OrbitPO: Learning Interface Equivariance for Tool-Using Language Agents

Abstract

Tool-using language agents act through software interfaces that are routinely refactored—tools renamed, fields reordered, entity handles relabeled—while the underlying tasks remain unchanged. Because such exact interface changes preserve task semantics yet still break capable agents, robustness to them is a prerequisite for integrating agents into evolving software. First, existing consistency objectives compare candidate-normalized action distributions, which conceal cross-interface shifts in unconditional action probability and yield no signal when only one can- didate is present. Second, view-local reward normalization and prefix matching discard cross-interface supervision precisely when equivalent rollouts fail or di- verge, a loss that grows as episodes lengthen. To address the first challenge, we pro- pose Orbit-Equivariant Policy Optimization (O RBIT PO), whose mass-preserving semantic alignment matches inverse-mapped unconditional candidate probabili- ties together with an explicit residual mass, supported by an exact information decomposition and non-asymptotic bounds that link retained-mass dispersion to alignment. To address the second, O RBIT PO introduces orbit-relative credit and counterfactual complete-history replay, which pool rewards across equivalent in- terfaces and align every visited history without requiring matched prefixes or successful continuations, with a coverage-dependent bound from replay loss to cross-interface return. Experiments on six core benchmarks, four frozen-transfer suites, nine adapted backbones, and controlled OrbitLab programs show that O R - BIT PO improves worst-view aggregate over exposure-matched augmentation on every backbone, by 6.7–9.9 points while keeping one policy call per decision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.