Position: Scientific Memory Is Not Chat History: Reframing Long-Term Memory Evaluation for Research Agents
Abstract
Long-term memory evaluation for language agents primarily measures preservation and retrieval of prior information across extended interaction horizons, but not how information changes across multi-session scientific interaction. This position paper argues for scientific memory as a distinct evaluation target since scientific interaction depends on preserving evolving information states under hypotheses, evidence provenance, contradiction, method and condition scope, revision, synthesis, and uncertainty as new information appears over time. Correctness is therefore defined by the evolving information-state transition Et−1 → Et rather than conversational consistency or retrieval alone. Prior benchmarks evaluate important components of scientific memory, yet do not evaluate whether an agent maintains one evolving information state across multiple sessions. We therefore argue that scientific-memory evaluation must distinguish retrieval failure, memory-update failure, and reasoning/provenance failure rather than collapse behavior into final-answer accuracy. To operationalize this position, we present a benchmark-oriented taxonomy of scientific-memory capabilities and failure modes and a five-condition diagnostic protocol (5DP) spanning online free retrieval, offline free retrieval, gold-source retrieval, gold-memory-state evaluation, and gold-evidence reasoning. Across 160 items and 5 systems, the model judge found that 19.8% of responses passing selected source, grounding, and fully-correct-answer checks failed joint informationstate requirements (95% interval [15.1, 24.7]), while the smaller human-calibration sample yielded a lower and uncertain gap estimate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.