acceptodds
Under review as a conference paper at ICLR 2027

The Context Gathering Decision Process: A POMDP Framework for Agentic Search

Abstract

A coding assistant asked about an unfamiliar repository cannot fit it into its prompt, so it searches: it issues a query, reads the result, and issues another. Harnesses run this phase, context gathering, by appending each passage to the prompt. One LLM call then chooses the next query, remembers what earlier rounds found, and decides when to stop. None of the three can be changed separately. We formalize the phase as a Context Gathering Decision Process, a finite-horizon POMDP over a fixed corpus, which separates two decisions the fused call makes together. The first is whether the agent's notes are rewritten into a bounded form after each observation, or accumulate. The second is who rewrites them: the policy, prompted to summarize itself, or the orchestrator, the code that runs between LLM calls. We call the latter an Orchestrator-Managed State (OMS). Rewriting raises accuracy over four published harnesses and three question-answering domains. Ownership does not change accuracy. It changes whether the agent declines a question its corpus cannot answer, where correct abstention rises points under OMS. A stopping rule reading the state in code ends stagnant searches early, with no LLM call. The gain from rewriting does not hold on two frontier policies. Accuracy and abstention therefore come from different parts of a harness, a reason to expose the update and the stopping rule as orchestrator interfaces. More broadly, pieces of an agent loop usually written as prompt engineering can be separated, measured and replaced without retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.