PANINI-DR: Controlled Deep Research Through Generative Semantic Workspaces
Abstract
Deep research requires connecting multiple steps of search, reasoning, and computation, with intermediate findings guiding what to investigate next. As investigations grow, agents must maintain research state across model calls and check findings before errors propagate into subsequent work. We introduce Panini-DR, a multi-agent harness that separates coordination of the overall investigation from the research and verification of individual subquestions. A shared workspace, built on the Generative Semantic Workspace representation, records subquestions, their dependencies, findings, and supporting evidence. Researchers investigate individual subquestions, while a separate verification agent checks proposed answers against those subquestions and their evidence. Verification feedback guides further investigation, and recorded findings inform subsequent assignments. We evaluate Panini-DR on BrowseComp-Plus, FRAMES, and MoNaCo against seven baseline implementations using a shared language-model backbone and frozen corpora. Panini-DR achieves the strongest average performance across the three benchmarks, with strong gains on harder questions. On BrowseComp-Plus, Panini-DR achieves **73.3%** accuracy, outperforming the next-best baseline by **16.0** percentage points, at higher inference cost. Its advantage is strongest on harder investigations, correctly answering **26 of 50** hard questions—nearly twice as many as the next-best baseline. Our ablation studies further show that structured planning and persistent state improve performance, as does verification combined with evidence checks. To assess whether correct answers are supported by the research performed, we also analyze the evidence retained by each system. Panini-DR achieves the highest answer recoverability from saved evidence on all three benchmarks, while complementary claim-support audits show stronger evidence-backed answers. Accounting for this evidence changes how systems compare and further strengthens Panini-DR's advantage over existing approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.