acceptodds
Under review as a conference paper at ICLR 2027

ContextWeave: A Real-World Workflow Benchmark

Abstract

Memory evaluation requires benchmarks that capture the temporal structure and cross-task dependencies of real work. We introduce ContextWeave, a longitudinal benchmark comprising 1,005 executable tasks reconstructed from 14 participants’ work records collected over multiple months, including 568 core evaluation tasks. It preserves chronological order, cross-task dependencies, and background tasks while protecting participant privacy. We compare performance on the same target tasks with and without recall under fixed histories and matched initial workspace states. Workspace Score measures task completion and output quality, while Preference Score measures adherence to participant-specific work practices. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both scores across all five tested base models. Experience-rich configurations outperform compact summaries and reduce repeated exploration, but also exhibit higher rates of memory-induced execution problems. These findings highlight the need to evaluate memory through downstream outcomes, behavioral effects, and robustness in sustained workflows.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.