acceptodds
Under review as a conference paper at ICLR 2027

PACT: Privileged Trace Co-Training for Multi-Turn Tool-Use Agents

Abstract

Multi-turn tool-use agents must reason, call tools, and adapt to observations across several interaction turns. Post-training such agents is challenging, as reinforcement learning often suffers from sparse rewards and weak credit assignment despite matching the prompt-only inference setting, while supervised fine-tuning on expert traces provides dense process supervision but can over-constrain the model to fixed trajectories. To tackle this, we propose PACT, a Privileged trAce Co-Training framework for multi-turn tool-use agents. The key idea is to use expert traces only as training-time optimization signals rather than rollout-time hints. PACT keeps rollout generation prompt-only, then uses expert traces to guide optimization through two complementary signals: a trace-conditioned RL surrogate that evaluates prompt-only rollouts under expert-trace context, and a component-aware SFT loss that provides annealed process supervision over reasoning and tool-call components. To reduce over-reliance on the training-only trace context, PACT incorporates prompt-only anchoring with standard prompt-only RL updates. Experiments on FTRL, BFCL, and ToolHop show that PACT consistently outperforms strong SFT- and RL-based baselines. Further analyses support the complementarity of the two training objectives and the benefits of component-aware supervision and prompt-only anchoring. These results highlight the effectiveness of privileged trace co-training for multi-turn tool-use learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.