acceptodds
Under review as a conference paper at ICLR 2027

Bayesian-Harness: Probabilistic Knowledge Control for LLM Agents

Abstract

An agent working on long-term tasks usually serialises its knowledge of, e.g., tools, interfaces, and codebases in its local environment. However, knowledge becomes stale and expires as the internal or external environment changes. Many existing harnesses update memory records after encountering an error, without explicitly modelling how knowledge expires over time. In this paper, we there- fore model knowledge staleness as a generative process from a Bayesian perspec- tive. We maintain two latent random variables for each memory, validity and quality. The validity is a binary latent state that fails at a Markov-modulated haz- ard and is restored only by a rewrite. We model quality in Bayesian form with a Beta-Bernoulli model over the memory’s success rate conditional on validity, with memories sharing statistical strength within each subsystem pool. We fur- ther build two datasets, MASKEDMIG and ACMESDK. MASKEDMIG extracts 157 real API changes from 18 Python libraries, each removing an old API and adding a named replacement in a release between April 2025 and August 2026. ACMESDK is a company Software Development Kit (SDK) benchmark. It is based on SOP- Bench and generated by Claude Code (Fable 5.1). Through extensive empirical investigation, from a 4B model to 27B and two closed-weight Gemini models, Bayesian-Harness is the best maintenance policy on every model of MASKEDMIG, 3.1–8.3 points above the strongest baseline of each, and in all 15 model–budget configurations of ACMESDK.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.