DataMemo: Turning Acquired Web Data into Reusable Agent Memory
Abstract
Web analytical tasks require agents to collect, join, and compute over records dispersed across web sources. When successive requests concern the same source, reacquiring those records wastes time and tokens, while retaining them only in unstructured interaction history makes coverage and interpretation difficult to track. Despite these challenges, dedicated benchmarks and systematic studies of how agents should manage and reuse web data remain limited. To evaluate this setting, we introduce DataMemoBench, a benchmark comprising 3,799 questions across eight domains to evaluate agents on successive analytical requests over shared web sources. We also present DataMemo, a relational agentic data memory framework that enables agents to incrementally store extracted records in a relational database, reuse them across questions, and acquire additional data as needed. Our evaluation shows that a continued-session baseline retains prior dialogue and files without explicitly representing the coverage and interpretation of retained data, while episodic and procedural memory still incur substantial overhead from repeated data acquisition and processing. With GPT-5.5 and Claude Opus 4.8, DataMemo reduces token cost by 36–46% and wall-clock time by 6–13%, while achieving a 5.8–11.8% relative improvement in accuracy over the strongest baseline, with efficiency gains growing as more questions reuse the accumulated data. Together, DataMemoBench and DataMemo provide an evaluation testbed and an effective framework for accurate and efficient web analysis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.