acceptodds
Under review as a conference paper at ICLR 2027

Pairwise Workflow Policy Optimization for Tool-Using Language Agents

Abstract

Tool-using language agents can automate complex tasks, but reinforcement learning treats different execution orders as distinct experiences even when they produce identical outcomes, introducing noise unrelated to task quality. We introduce Pairwise Workflow Policy Optimization (PWPO), which verifies that swapping two adjacent tool calls preserves their outputs and resulting state, then optimizes the summed probabilities of both complete execution histories. In deterministic environments with restorable states, this aggregation preserves the expected on-policy gradient and removes variation within verified pairs; our analysis specifies when this reduction offsets additional computation. Across AppWorld-Atomic, ALFWorld, and WebShop under matched rollout budgets, PWPO achieves the highest means among evaluated methods on all 16 main metrics, improving AppWorld-Atomic task completion over G2PO by 2.32–2.64 percentage points. These results support learning from verified workflow alternatives while retaining the agent's sequential deployment interface.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.