SciEngram: Evidence-Grounded Scientific Memory for Cumulative Research Reasoning
Abstract
Scientific agents that continually read the literature must maintain traceable and revisable scientific beliefs. Existing scientific question answering, retrieval-augmented generation, knowledge graphs, and general-purpose agent memory mainly evaluate query-time retrieval and answer correctness. They do not test whether the evidence structure behind an answer persists across queries and changes with new evidence. We introduce SciEngram, which formalizes scientific memory as an evaluable and implementable computational problem. SciEngram-Benchmark evaluates incremental scientific question answering through 401 questions in six task families, answered as papers arrive within literature groups. An ordered-evidence audit design specifies complementary revision and retention checks. It evaluates answers along eight dimensions, including correctness, evidence quality, source traceability, relational reasoning, scope qualification, and abstention. SciEngram-Core implements these functions with provenance-bound paper memory, scope-aware consolidation, dependency correction, and factorized belief states. Under the main reader configuration, SciEngram-Core scores 8.87, exceeding the strongest BM25 baseline at 8.65, and improves source traceability by 1.09 points. Removing hierarchical evidence organization or dependency correction lowers the score by more than one point. The gains therefore come primarily from sustained organization of evidence sources, relations, and scope.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.