OWG: Structuring Agent Evidence as Outcome-Witness Graphs for Efficient Long-Horizon Task Closure
Abstract
An agent can retain its entire interaction history and still lose track of what remains to be done. Long transcripts place decisive evidence alongside obsolete outputs, repeated observations, and failures unrelated to the requested result. We introduce OWG, a training free method that organizes an agent's context around the evidence for completing its current task. An Outcome-Witness Graph links each requirement to observations that support or contradict it, the output being produced, and the checks performed on that output. After every tool response, OWG updates these relations and uses the same state to select the next prompt, direct the next action, and decide when to return the result. Raw observations remain available for targeted recovery, while completed work and repeated messages leave the active context. Across six benchmarks and five models, OWG lowers cost in all 30 evaluated settings, with task success increasing in 17, remaining unchanged in six, and decreasing in seven. Five runs on fixed PinchBench and SWE Bench subsets reduce tokens by 53.2% and 74.2%, respectively, while Pass@1 increases by 0.4 percentage points on each. These results identify task evidence as a useful basis for reducing the cost of sustained agent execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.