EDO: Expectation–Observation Residuals for Token-Efficient LLM Agents
Abstract
Large language model agents solve practical tasks through repeated tool interactions, but long tool outputs accumulate in their context and make subsequent inference increasingly expensive. Compressing these outputs requires more than retaining many relevant facts: feedback must also communicate how the actual result confirms or revises the agent's pre-execution expectations. We study this distinction between fact coverage and evidence updating, and introduce EDO, an expectation-conditioned representation of tool feedback. Before each tool call, the main agent writes a compact expected observation grounded in its current state. After execution, a smaller model compares the actual output with this expectation and constructs a semantic residual organized around what the result confirms, corrects, newly reveals, or leaves unresolved. The expectation remains an explicit reference rather than environmental evidence; for long outputs routed through model-based extraction, the expectation–residual pair replaces the raw output in subsequent context. Across six software-engineering and terminal-agent benchmarks, EDO achieves 98.97% of Raw's macro-averaged task score while using, on average, 59.06% of its main-agent tokens and 60.79% of its estimated cost after benchmark-wise normalization. On ProgramBench, it exceeds a matched summary by 4.10 percentage points; offline diagnostics show that it raises correction recall from 70.00% to 94.55% without increasing overall fact recall.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.