acceptodds
Under review as a conference paper at ICLR 2027

How Many Parameters Is a Trace Worth? Measuring Borrowed Capability in LLM Agents

Abstract

Language-model agents rarely operate in a vacuum, as they typically have access to previous requests and execution traces. These traces are especially valuable for systems that rely on small local models, where traces produced by a larger model can guide a smaller, faster one. In this paper, we study borrowed capability: the degree to which a weaker model, given a stronger model's verified execution trace on a related task, can perform as if it had the stronger model's capability. We derive 1,200 targets from 300 source tasks across six agentic benchmarks; each target changes the source's arguments, context, procedure, or goal, so that replaying the source trace fails. Across 20 open checkpoints from four families (270M-34B), each attempting every target with and without the trace, we find that a trace is worth roughly doubling the executor's size, provided that the executor is competent in the environment, i.e., solves at least one of its original tasks on its own. For competent executors, trace access raises exact success by 12 points, about as much as replacing the executor with a same-family model twice its size. Within an environment, larger executors close a larger share of the gap to the trace's author, although the smallest ones gain the most points. Executors below this competence threshold gain only 1.5 points, unless what they fail is the required answer format, which the trace demonstrates. Traces help least when the interface has changed, and there they can mislead. A trace, in short, can make a small agent act like a larger one, but it cannot teach an agent to act.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.