HINT: Learning the Harness Protocol Interface-Invariant Training for SWE Agents
Abstract
Large language model (LLM) agents now resolve real GitHub issues by exploring a repository, editing files and running tests. A model does this through a harness such as Claude Code, Codex CLI, OpenCode or OpenHands SDK: the scaffold that decides which tools it sees, what they are called, and how calls and results are formatted. These details differ across harnesses and they matter. With model and tasks fixed, changing only the harness moves SWE-bench Verified accuracy by up to 9.0 points, and fine-tuning on trajectories from one harness teaches the interface along with the task: the gain is largest on the source harness and absent on another, widening the spread across harnesses from 3.4 to 11.7 points. Examining 578,856 real tool calls from four production harnesses, we find that they all reduce to the same 12 semantic actions. We formalize this shared core as the Harness Protocol and use it to separate a harness into two layers: the protocol trajectory, the actions and observations every harness must carry, and a surface of tool names, documentation, output and turn formatting. Rendering a protocol trajectory into any surface and parsing it back recovers the same actions (1.87M checks, no mismatches). HINT (Harness-INvariant Training) exploits this: it parses each real trajectory into the protocol and re-renders it under many sampled surfaces, some with meaningless tool names, so that one harness's data trains an agent for any harness with no new task data. Fine-tuning Qwen3.6-35B-A3B on Claude Code trajectories, HINT matches a step-matched control on Claude Code itself, leads under every unseen production harness by up to 8.7 points, narrows the SWE-bench Multilingual accuracy range across five harnesses from 7.4 to 1.6 points, and is on a par with training on real trajectories from four harnesses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.