acceptodds
Under review as a conference paper at ICLR 2027

BranchWise: A Path Is Not a Policy — Unfolding Successful Trajectories into Model-Tailored Decision Boundaries

Abstract

Tool-use agents make sequential decisions among multiple possible actions, yet successful trajectories expose only one realized path. Direct supervised fine-tuning (SFT) underutilizes their latent supervision: long trajectories entangle local decisions and expose only one reference branch per state, leaving alternative failure directions implicit; static demonstrations do not explicitly reveal model-specific errors; and uniform token-level supervision does not explicitly target the regions that distinguish behaviors. When trained on general tool-use data, SFT can therefore transfer inconsistently and even underperform the base model. We propose BranchWise, a post-training framework that unfolds successful trajectories into model-tailored local decision boundaries. It decomposes trajectories into state-conditioned decisions and constructs verified counterfactuals from tool constraints and execution evidence. These counterfactuals are fused with verifiable error branches generated by probing the initial policy to form compact local decision boundaries. BranchWise localizes comparisons to decision forks and combines evidence-calibrated reference-branch supervision with margin-aware listwise ranking to emphasize insufficiently separated boundaries. With all methods trained from the same general tool-use corpus rather than benchmark-specific supervision, BranchWise outperforms Vanilla SFT across -Bench, BFCL V3, ToolHop, and ToolTalk-Hard. It achieves the best overall performance among all compared methods, demonstrating stronger and more consistent downstream transfer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.