AutoScientist: An Autonomous Agent for Long-Horizon Scientific Discovery via Connectome Memory
Abstract
Accelerating the rate of scientific discovery requires autonomous agents capable of continuous literature assimilation, rigorous hypothesis generation, and empirical verification over extended horizons. Existing autonomous research agents rely on linear flat-file scratchpads or unindexed vector databases for memory. These representations suffer from context window saturation, catastrophic forgetting of early observations, and an inability to detect contradictions across multihop causal chains. To resolve these limitations, we introduce AutoScientist, an autonomous discovery system organized around a biologically inspired, multirelational graph connectome. AutoScientist treats memory as an active graph, modelling scientific knowledge as an evolving network of typed concepts, causal relations, and empirical findings whose structure persists as it grows. The architecture executes diurnal discovery cycles: active phases acquire literature, formulate structured hypotheses, and run computational simulations in code sandboxes, while nocturnal phases apply multi-pass consolidation algorithms that resolve logical contradictions, prune unreinforced associations, and synthesize structural analogies. We evaluate AutoScientist across 20 longitudinal discovery cycles and three disparate scientific domains: non-equilibrium statistical mechanics, systems biology gene regulatory circuits, and non-binary error-correcting codes. Empirical evaluations show that the agent’s memory develops a sparse, modular small-world topology that no degree-matched random null reproduces. In controlled multihop retrieval and falsification benchmarks, graph memory holds every chain step and resolves every recorded contradiction, whereas dense vector retrieval recovers only a third of a chain’s interior, never exceeds chance on a recorded contradiction, and fills nearly half of its retrieved context with decoys.A sweep of workingmemory capacity shows that bounded focus strikes an optimal balance between prompt efficiency and thrashing avoidance, and neuromodulatory state transitions prevent stagnation on intractable problems. The full agent runs on a single locally hosted 32B open-weight model, and matches the suite scores reported for frontier commercial systems on the same long-memory sets, so strong long-horizon scientific discovery is within reach of modest, fully local hardware. These results establish connectome-based memory as an effective computational foundation for long-horizon autonomous scientific research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.