ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness
Abstract
Large language model (LLM) agents increasingly operate within a harness, the scaffolding that determines what enters the executor’s context, yet the experi- ence they accumulate across tasks rarely flows back into this harness. Existing approaches include executor fine-tuning and external memory retrieval, but combin- ing task-adaptive retrieval with experience reuse across frozen executors remains challenging. To this end, we introduce ExpHarness, a learnable experience harness that improves frozen and replaceable LLM executors without modifying their parameters. Specifically, ExpHarness distills trajectories into reusable skills and failure lessons within a self-evolving experience graph, and trains a lightweight retrieval copilot that decides, per task, how broadly to explore the graph and how strongly to favor historically useful experiences over merely similar ones. The copilot is optimized with reinforcement learning from a utility-grounded reward combining the with/without-experience score difference and an absolute- performance term; the same reward updates the graph during training. Extensive experiments on ExpSuite, spanning 10 static benchmarks and 2 agentic environ- ments, show that ExpHarness improves over the strongest baseline by 12.1% and 4.5% on static tasks and by 21.4% and 12.7% on agentic tasks with the smaller and larger executors, respectively, while reducing interaction steps by up to 21.6%. Transfer experiments further examine reuse of the learned harness across executors of different scales and reasoning capabilities, with joint graph and copilot transfer performing closest to target-specific training among the evaluated transfer variants.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.