Reverie: An Inference-Time Cognitive Architecture for Continuous LLM Deliberation
Abstract
LLMs are reactive: a prompt arrives, the model computes, and cognition halts until the next prompt. Existing multi-agent LLM systems inherit this constraint: they coordinate multiple calls per response but still respond only when prompted. We introduce Reverie, an inference-time architecture in which a set of specialised agents runs continuously on a shared workspace where ideas accumulate, decay, and compete for influence. An executive composes a response when the user speaks, or on its own when accumulated internal activity crosses a threshold; between turns, the system keeps thinking. On a downstream task of research assistance over open-ended scientific questions, five independent human raters, blinded to system identity, preferred Reverie's conversations over Mixture-of-Agents on pairs (; mixed-effects CI –, ), with every rater preferring Reverie. At matched compute, Reverie also produces more conceptually distinct thoughts per turn than Mixture-of-Agents and more than Self-Consistency (Vendi Score, Bonferroni-significant), at roughly lower cloud-equivalent cost per session than Mixture-of Agents. On nine verifiable benchmarks (factual QA, math, code) Reverie is statistically tied with Mixture-of-Agents on eight and worse on rare-fact retrieval (PopQA), so the cost of continuous deliberation on convergent tasks is small.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.