ContextWeave: A Real-World Workflow Benchmark
Abstract
Memory evaluation requires benchmarks that capture the temporal structure and cross-task dependencies of real work. We introduce ContextWeave, a longitudinal benchmark comprising 1,005 executable tasks reconstructed from 14 participants’ work records collected over multiple months, including 568 core evaluation tasks. It preserves chronological order, cross-task dependencies, and background tasks while protecting participant privacy. We compare performance on the same target tasks with and without recall under fixed histories and matched initial workspace states. Workspace Score measures task completion and output quality, while Preference Score measures adherence to participant-specific work practices. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both scores across all five tested base models. Experience-rich configurations outperform compact summaries and reduce repeated exploration, but also exhibit higher rates of memory-induced execution problems. These findings highlight the need to evaluate memory through downstream outcomes, behavioral effects, and robustness in sustained workflows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.