acceptodds
Under review as a conference paper at ICLR 2027

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

Abstract

For language-model agents, outcome-based reinforcement learning provides verified task feedback but assigns the same trajectory-level advantage to every intermediate decision, making it difficult to identify which actions contributed to success or failure. Process supervision can refine this credit, yet treating it as an independent auxiliary objective may encourage or discourage intermediate decisions independently of the trajectory’s outcome advantage. Self-distillation with privileged information provides finer-grained guidance through changes in response likelihood. However, these changes do not directly specify how to allocate outcome credit, and the guidance may depend on conditions absent from the target trajectory. We introduce TASPO, a framework for constrained credit redistribution. Outcome feedback sets the trajectory-level advantage; process supervision redistributes it across decisions while preserving its sign and mean within each trajectory. TASPO aligns guidance from verified successful sibling trajectories with the completed target interaction, measures guidance-induced support over executable actions, and converts these scores into positive, bounded, mean-one weights. The constraints preserve advantage polarity and decision-level mean while limiting local reweighting, without additional environment interaction. Across ALFWorld, Search-QA, and WebShop with three backbone models, TASPO improves over GRPO by 12.4 percentage points on average. Analyses confirm that both alignment and action-level aggregation are essential, while TASPO stabilizes policy-KL optimization dynamics. Codes are available at https://anonymous.4open.science/r/TASPO-0DAA/https://anonymous.4open.science/r/TASPO-0DAA/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.