Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives
Abstract
Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. Existing approaches remain brittle in multi-step and multi-turn settings, where the model must map one tool's output into the next tool's input schema, and degrade further as the tool catalogue grows. We introduce Tool Primitives, which replace schema-based invocation with a natural language interface: each tool is wrapped with an LLM that resolves its own schema and executes the underlying function, so callers describe what they want instead of emitting schema-compliant arguments, and dependent calls pass results to one another in natural language. Building on them, we host ToolFace, a repository of 25,519 functions from which only the tools retrieved for the current query enter context, bounding what the model reasons over by the retrieval budget rather than by the size of the repository. A primitive, however, answers only the request it is handed. We therefore propose HEART, a Harness Engineering framework via Agent-native, Reusable Tool Primitives, which decomposes a query into an invocation plan, grounds each step's arguments from the query and from earlier results, and validates every returned result, feeding failures back as targeted revisions. Experiments on five benchmarks demonstrate that HEART outperforms SFT-based models by 10% on average and surpasses GPT-5.4, Claude-4.6-Sonnet, and Gemini-3.1-Pro by 6% on average while reducing API cost by up to 85%. On 50 real-world tasks, HEART achieves 84% task completion, the average of three frontier commercial models (22%).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.