Stored But Not Delivered: Benchmarking and Repairing Research-State Discontinuity in Scientific Agents
Abstract
In scientific agents, each hypothesis, experiment, and result conditions later decisions, and a decision can depend on evidence produced many steps earlier. An agent can store a hypothesis and the result that refutes it, yet retrieve one without the other, and so repeat a direction it has already ruled out. We formulate this as research-state discontinuity: the context delivered for a decision can omit relationships that stored history still holds, and inspecting the store or final task success does not reveal the omission. To assess it we construct the Executable Scientific-Agent Memory Benchmark (ESMB), a controlled diagnostic suite that fixes histories, checkpoints and authored action effects so that only memory varies, and scores a memory interface returning labeled evidence separately from a planner acting on it. This separation localizes the failure to two operations. Text baselines confine candidates to the current branch, so the decisive record is dropped before ranking, and a record that does reach ranking can still be displaced once the budget fills. We then propose Research State Memory (RSM), which gives every item a branch identity, stores contradiction and supersession as explicit edges, and delivers an item together with the evidence bearing on it. RSM raises joint successes from 6/15 to 12/15 on the cued suite, and solves 60/102 externally authored ScienceAgentBench tasks against repair summary’s 46/102.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.