DyadMem: Dual-Zone Memory with Source-Grounded Recall for Long-Horizon Agent Tasks
Abstract
Long-horizon agentic tasks produce trajectories that exceed LLM context windows, creating a compression–traceability dilemma: compact memories support global understanding but can obscure details and provenance, whereas chunk retrieval does not explicitly preserve trajectory structure. Agent histories therefore require both an up-to-date task state and a way to locate precise evidence across ordered action–observation steps. We propose DyadMem, a dual-zone memory framework comprising Situation Memory and Event List. Situation Memory maintains a cumulative summary of goals, progress, decisions, and constraints, while Event List forms a chronological index whose stable identifiers map to immutably preserved raw records. Memory construction is question-independent, and hierarchical re-compression preserves identifier coverage. At query time, the agent drafts an answer from memory and optionally verifies it through find_raw_events, which supports natural-language search, identifier-based reading, and focused re-reading. On CAME-Bench, DyadMem achieves 76.17% F1 with Gemma-4-31B and 64.29% with Qwen3-32B; on AMA-Bench, it reaches 67.75% and 66.15% accuracy, respectively, establishing new state-of-the-art results across both benchmarks and outperforming eight strong baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.