Quality Meets Structure: Context-Aware Trajectory Selection for Agentic Supervised Fine-Tuning
Abstract
The effectiveness of agentic supervised fine-tuning (SFT) depends on the quality of training trajectories and the diversity of their interaction patterns. However, existing data selection methods often overlook how decisions depend on earlier interactions and how execution patterns overlap across trajectories, resulting in redundant or less informative supervision. We introduce CATS, a graph-based selection method that estimates trajectory quality from context dependence and action complexity by comparing a frozen small language model's prediction losses across context variants. It constructs a similarity graph from ordered sequences of tool calls, natural-language replies, and tool-return statuses to capture overlap in execution patterns. Through message passing, CATS aggregates similarity-weighted trajectory scores, iteratively selects the highest-priority trajectory, and updates the priorities of unselected neighbors to discourage structural redundancy. Selecting 100K trajectories from a pool of over 600K, CATS outperforms nine baselines in average agentic and general performance across four benchmarks, improving average agentic performance by up to 2.6 percentage points over the strongest baseline. Ablations demonstrate the contributions of both quality scoring and structural selection, while subset analyses show broader coverage of execution patterns with high average trajectory quality. These findings highlight the value of jointly considering context-aware decisions and execution-pattern diversity for effective data selection in agentic SFT.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.