acceptodds
Under review as a conference paper at ICLR 2027

Beyond Outcome Filtering: Diagnosing and Correcting Behavioral Bias in Agentic SFT

Abstract

Supervised fine-tuning (SFT) for agentic language models is often built from demonstrations selected by task outcome. This creates a mismatch: curation evaluates whether a trajectory succeeds, while SFT learns from the entire trajectory. Successful demonstrations can therefore retain undesirable behaviors such as redundant retrieval, repetition, excessive verbosity, and failures to use available evidence. We introduce DiaCurate (Diagnose-and-Curate), a diagnostic-driven framework for agentic SFT data curation that identifies behavioral failures, applies targeted edits to demonstrations, and re-diagnoses the policy after retraining to detect collateral regressions. On a deep-research agent evaluated with HLE and BrowseComp, SFT on an outcome-filtered reference corpus improves task accuracy while increasing search inefficiency by 25.9 and 17.2 points and verbosity by 19.2 and 26.1 points on corresponding development sets. Directly shortening trajectories reduces unnecessary search but increases premature commitment and degrades accuracy. Re-diagnosis motivates an evidence-use refinement that reduces search inefficiency while preserving task performance, reaching 35.2 and 34.1 avg@3 versus 33.5 and 33.3 for the reference. These results show that successful trajectories are not necessarily good demonstrations: agentic SFT curation should account for both task outcomes and the behaviors learned from training trajectories.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.