acceptodds
Under review as a conference paper at ICLR 2027

Unlocking Latent Tool-Use Capabilities with Procedural Hint Supervision

Abstract

Reinforcement learning with outcome-based rewards can improve tool use, but a model may know how to execute individual actions without reliably organizing them into successful procedures. Sparse task-level feedback provides little guidance on how to correct these procedural failures. We introduce Procedural Hint Supervision (PHS), which uses reusable guidance to improve credit assignment and help models apply their existing capabilities more reliably. PHS extracts procedures by comparing successful attempts from a more capable model with failed attempts from the model being trained. During training, selected guidance is used to re-score the model's own trajectories, and differences between hinted and unhinted token log probabilities shape token-level advantages in two extended formulations of Group Relative Policy Optimization (GRPO). The trained policies are evaluated without hints on held-out sets. Trained separately on AppWorld and EnterpriseOps, both formulations improve held-out success and reciprocal transfer, outperforming vanilla GRPO. Across four seeds, the largest pass@4 gains over GRPO are 19.1 percentage points on AppWorld and 7.4 percentage points on EnterpriseOps. The same AppWorld policy gains 14.9 percentage points on its harder test-challenge split. Without targeted benchmark training, gains over GRPO reach 4.1 percentage points in BFCL multi-turn pass@4 for an EnterpriseOps-trained policy and 10.4 percentage points in mean pass across retail, airline, and telecom for an AppWorld-trained policy. Together, the in-domain test sets and transfer results suggest that procedural guidance can unlock latent capabilities, turning existing tool-use skills into reliable behavior that generalizes across tasks, tools, and environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.