The Context Trade-Off: Understanding Information Use in Language Models and Agents
Abstract
Large language models face a fundamental tension at inference time: additional context provides more evidence, yet longer inputs can induce context rot—the degradation of model performance as input length grows. The same tension arises in language model agents as tool responses accumulate in their history. We ask when additional context helps or hurts, and how context management changes this trade-off. Across multiple benchmarks, evaluation formats, and model families, performance often rises and then falls as context grows. We call this pattern context saturation. Its shape and peak vary across models and tasks. To explain this pattern, we compare a model with an ideal solver: a reader that sees the same input and extracts everything it implies about the task, so its performance measures how much information the input contains. Two properties make this comparison useful. Adding background text unrelated to the answer cannot lower the ideal solver's performance, and neither can rewriting an agent's history, as long as earlier observations can still be retrieved. Under such changes, any drop in the model's performance is therefore a failure to use information that is still available, and any gain is information made easier to use. Controlled experiments across benchmarks and model families support three findings. (1) In most controlled settings, the rise reflects evidence arriving and the fall reflects evidence becoming harder to use, so the peak depends on the model. (2) Performance falls even when all needed evidence is kept. Evidence is also used jointly: a piece of evidence without its complement is lost more often, so a missing piece can matter more in longer contexts. (3) For agents, rewriting the same observations into a compact form can raise success several-fold. Rewritten histories fail when the agent reads past its target or can no longer revisit earlier observations. In these controlled settings, long context fails not for lack of information but because information becomes hard to use.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.