HoMe: Toward Practical Harness Optimization of Memory Systems for AI Assistants
Abstract
The behavior of memory systems for long-running AI assistants is governed by a harness, which consists of the code, prompts, and configurations for constructing, maintaining, and retrieving memory. However, optimizing the harness of a memory system remains largely manual, and existing automated methods restrict either the editable components or the search procedure. In this paper, we introduce HoMe, a pure evaluation-driven framework that iteratively evaluates candidate harnesses and derives each new candidate from any previously evaluated harness by editing any part of it. HoMe combines session-level memory versioning with execution traces to localize failures to the responsible sessions and harness components. It further reduces evaluation cost through smoke tests, candidate gating, adaptive sampling, and the reuse of constructed memory. We validate HoMe with the open-source memory system on two large-scale long-horizon memory benchmarks, LongMemEval-S and BEAM. With only six evaluated candidates per benchmark, HoMe discovers harnesses that improve end-task performance over the originally strong harness while consuming far fewer query-time tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.