acceptodds
Under review as a conference paper at ICLR 2027

Where Context Goes and What It Owes: Residency Contracts for Long-Running LLM Agents

Abstract

Long-running LLM agents compact histories, delegate subtasks, prune tool outputs, and resume sessions. These transformations can preserve task information while losing its conditions of use, such as exact wording, scope, source role, retrievability, freshness, and supporting evidence. Context management must therefore determine what survives each boundary, in what form, for a particular receiver. We introduce **residency contracts**, which make these requirements explicit for each typed artifact. **Boundary replay** restores a pre-boundary checkpoint, perturbs one condition of use, and compares paired continuations to measure its signed behavioral effect. Replay supplies offline evidence and supervision, while a runtime controller combines policy rules with a learned scorer to produce contracts without receiver-time replay. On BoundaryBench, our benchmark of constructed coding tasks, twelve executor models share 24 task templates. The controller improves preservation by 0.20 to 0.22 and strict success by 20.8 to 25.0 percentage points over their native-harness baselines. For Grok 4.20 non-reasoning, preservation rises from 0.61 to 0.83 and success from 0.30 to 0.55 while input token use falls by 55%. Relevance-only scoring reduces preservation by 0.09. These paired comparisons evaluate the combined scorer-and-rule policy with construction metadata under different token caps. The results support explicit conditions of use as an object for context management on the evaluated families.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.