acceptodds
Under review as a conference paper at ICLR 2027

APPO: Agentic Procedural Policy Optimization

Abstract

Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit over coarse heuristic units, such as tool-call boundaries or fixed workflows, making it difficult to identify which intermediate decisions influence downstream outcomes. In this work, we study agentic RL from two perspectives: where to branch and how to assign credit after branching. Our pilot analysis shows that high-entropy candidate positions are distributed throughout generated reasoning rather than concentrated only at tool calls, while entropy alone does not reliably order observed branch outcomes. Motivated by these observations, we propose Agentic Procedural Policy Optimization (APPO), which shifts branching and credit assignment from coarse interaction units to fine-grained decision points in the sequence. APPO selects branching locations using a future-aware Branching Score that combines token uncertainty with a discounted suffix policy-shift signal. It further introduces dual-group advantage estimation and procedure-level advantage scaling to provide more targeted credit after branching. We evaluate APPO on 13 benchmarks spanning mathematical reasoning, knowledge-intensive reasoning, and deep search, and find consistent improvements over strong agentic RL baselines by nearly 4 points on aggregate, while retaining efficient tool use and interpretable behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.