acceptodds
Under review as a conference paper at ICLR 2027

EHR-Evo: Evolving Harnesses for EHR Agents through Execution Experience

Abstract

Large language model agents provide a promising interface for accessing electronic health records (EHRs), but reliable execution remains difficult because database queries require complex schema navigation, tool use, and multi-step program construction. Existing systems can recover from individual failures through retries and execution feedback, yet typically rely on fixed execution scaffolds and make limited use of accumulated experience. We introduce , an experience-driven execution harness around a frozen language-model agent. An outer-loop proposer uses sanitized execution traces to evolve model-external planning, retrieval, memory, and recovery mechanisms, after which the selected harness is frozen for evaluation. During execution, the harness retrieves reusable operational knowledge and solution traces, constructs question-specific guidance, and uses execution feedback for targeted recovery and cross-question experience reuse. We evaluate on MIMIC-III, eICU, and TREQS, comparing the complete evolved harness with fixed-retry, persistent-memory, and existing agent/harness baselines. The results show consistent improvements in answer success while revealing distinct effects of recovery and persistent experience on execution completion. Our study demonstrates that EHR-agent capability can be improved by evolving and adapting the model-external harness while keeping the underlying solver frozen.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.