ParSer: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Abstract
Reasoning over documents far beyond an LLM's context window remains challenging, as evidence may be sparsely distributed across hundreds of thousands of tokens. Sequential memory agents stream documents chunk by chunk into a compact recurrent memory, enabling bounded-context processing of arbitrarily long inputs. However, because every reading step immediately updates the memory on which subsequent processing depends, these agents couple document traversal with sequential reasoning. This coupling makes the reasoning sensitive to evidence position and forces the sequential inference path to grow with document length. We introduce **ParSer** (**Pa**rallel **R**eading, **Se**quential **R**easoning), a framework that separates document reading from question reasoning. To realize this separation, ParSer assigns local reading to chunk-bound subagents and global reasoning to a lead agent. The subagents read in parallel under the lead agent's queries; the lead agent integrates their findings and iteratively refines its queries as evidence accumulates. This decoupled design concentrates all learnable behavior in the lead agent and optimizes it with reinforcement learning, while the lightweight chunk readers remain frozen off-the-shelf models. On multi-hop QA with contexts ranging from 7K to 896K tokens, ParSer with a 4B backbone beats the strongest sequential memory baseline by 5.7 points on average and by 12.0 points at 896K tokens. With a 9B backbone, ParSer surpasses DeepSeek-V4-Pro by 6.3 points. Experiments show that ParSer remains robust to changes in evidence position, order, and distance that cause large accuracy swings in sequential methods. By parallelizing document coverage, ParSer reduces inference latency by up to .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.