acceptodds
Under review as a conference paper at ICLR 2027

Learning to Solve and Execute: Trajectory Refinement for Harness-Aware Agentic Fine-Tuning

Abstract

Language-model agents solve difficult real world tasks under tool execution environments and harness constraints. Supervised fine-tuning, usually called SFT, typically uses demonstrations from stronger teachers, yet these may not cover the states encountered by students. Revising student trajectories offers a way to address this mismatch, but changing an action can invalidate subsequent observations. We introduce Repair SFT, which constructs offline supervision by asking a black-box teacher to correct causal errors in student trajectories while retaining valid behavior. Revised tool calls are executed in the environment to regenerate the affected continuation, keeping actions and observations consistent. Hybrid SFT selects between repaired trajectories and direct teacher demonstrations using training-task outcomes and agent roles. We evaluate both methods with Qwen2.5-14B and Qwen3.5-9B under a shared memory harness. On the Qwen3.5 experiment, Repair SFT reduces step-budget exhaustion from 44.1% to 14.3%, while Hybrid SFT raises MuSiQue F1 from 20.24% to 49.59%. On Qwen2.5, Hybrid SFT improves MuSiQue and Qasper F1 over Direct SFT and achieves the highest IIRC F1 among the three SFT variants. These findings support student-conditioned trajectory repair and selective supervision for agentic SFT.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.