LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via State Proprioception
Abstract
Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Existing systems delegate context management to runtime rules or train agents to compress their working memory, often leaving runtime state implicit or discarding evidence through summarization. We argue that a basic obstacle is context-state blindness: agents must decide what to keep or archive without an explicit view of block costs and remaining capacity. We introduce VISTA (Visible Internal State for Tool Agents), a training-free, model-agnostic layer that represents working memory as typed, addressable blocks, exposes token usage, creation age, archive status, and remaining budget through a runtime dashboard, and archives blocks as recoverable full-fidelity payloads. Across LOCA-Bench, BrowseComp-Plus, and GAIA, the same untrained interface transfers across trajectory scales. On LOCA-Bench, it raises Gemini-3-Flash accuracy from 22.7% to 50.7%; it reaches 58.0% on BrowseComp-Plus and remains competitive on GAIA. Pressure sweeps characterize its operating range, gains transfer across four backbones, and ablations show that the dashboard contributes beyond archive and recovery tools.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.