OpCtx: Evolving Event Views for Budget-Constrained Agent Memory
Abstract
Long-term memory enables language agents to accumulate knowledge across interactions. However, as histories grow, answering new questions requires organizing observations of events that evolve across interactions and recovering useful evidence within limited reading budgets. These challenges are amplified in multimodal settings, where images contain overlapping details about objects, attributes, and relations. Existing memory pipelines still face two difficulties in addressing these challenges. First, they leave relationships among observations of the same event implicit, requiring the answering model to reconstruct these relationships from scattered evidence. Second, they retrieve multiple facts conveying the same information when selection relies primarily on similarity, leaving less of the retrieval budget for complementary evidence. Retrieving more candidates and applying model-based reranking can improve selection, but reranking costs grow with candidate text. To address these difficulties, we propose OpCtx, a memory architecture connecting online event organization with question-conditioned evidence reading. OpCtx organizes source-linked observations into cross-interaction event views and, using a training-free Reader, aggregates selected facts and supported event information into structured evidence. This helps the answering model interpret relationships and changes across interactions. To improve evidence coverage within the reading budget, a coverage-preserving selector balances lexical and semantic signals under explicit fact-count limits and an optional character budget. Experiments on MemEye, LoCoMo, and LongMemEval demonstrate the effectiveness of OpCtx across multimodal and text-only memory tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.