acceptodds
Under review as a conference paper at ICLR 2027

ScreenMem: Benchmarking Proactive Memory Maintenance over Long-Horizon Computer-Use Streams

Abstract

Memory is a fundamental capability of computer-use agents, enabling them to retain and reuse information accumulated throughout everyday computer use. Yet existing evaluations do not fully characterize a key setting for persistent memory: continuously observing long-horizon device activity, deciding what to preserve before future information needs are known, and later answering from the resulting memory without revisiting the original stream. We introduce ScreenMem, a benchmark for proactive, query-agnostic memory maintenance over day-long computer-use streams. ScreenMem reconstructs application interfaces and replays interactions from open-source real-world traces, producing 10 user-day streams across PC and mobile devices totaling 100.1 hours, together with 484 human-verified questions covering facts and states, state changes, activities and events, and behavioral patterns. Under a unified causal-ingestion protocol, we evaluate 15 existing memory systems spanning recent-context, implicit visual, explicit multimodal, and generic text-memory paradigms on both PC and mobile streams. Across 15 systems, the strongest evaluated system reaches only 31.4% overall accuracy, while State Evolution accuracy peaks at 15.8%. Further analyses reveal substantial variation across temporal horizons, together with pronounced differences in ingestion and storage efficiency. ScreenMem can facilitate systematic analysis of current memory systems and further research on persistent memory maintenance for computer-use agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.