When Coding Skills Bind to Their Harness: Understanding and Mitigating Cross-Harness Failures in Agentic SFT
Abstract
Agentic supervised fine-tuning (SFT) can improve an agent on the harness that generated its trajectories while degrading performance after deployment through another harness. We identify this failure in Qwen3.5-9B: SFT on Claude Code trajectories raises native success but lowers success on the held-out DeepSeek Harness by 9.8 percentage points. We call this failure harness entanglement. Controlled interventions reveal sensitivity to tool presentation and identifiers. We introduce ARCHER (Across-task Randomized Consistency for Harness Entanglement Reduction), a cross-task shuffled-pair objective that trains different harness renderings with position-level loss consistency. On SWE-bench Verified, ARCHER reaches 62.2% and 63.2% on the held-out DeepSeek Harness and Pi, respectively, compared with 50.6% and 49.9% for the base model, while matching the base model on Claude Code and improving OpenHands. Ablations favor the shuffled construction and show higher transfer scores with position-level consistency. These results show that native-harness scores can hide deployment failures and that randomized consistency can reduce this dependence during SFT.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.