acceptodds
Under review as a conference paper at ICLR 2027

daVinci-Searcher: Eliciting Deep Research Capabilities Data Efficiently

Abstract

Deep research agents are increasingly important for solving complex, information-intensive tasks that require sustained search and evidence integration, and training on synthetic trajectories has enabled open-source agents to achieve performance competitive with proprietary systems. We ask: *Can a small set of high-quality trajectories elicit strong deep research capabilities ***data efficiently*** through supervised fine-tuning alone?* We introduce **daVinci-Searcher**, which combines scalable task synthesis with trajectory selection. Our Wikidata-based pipeline constructs challenging, verifiable multi-entity questions, using obfuscated entities and cross-entity constraints to control difficulty and encourage sustained search. We further identify *loop-prone reasoning* in teacher trajectories, where excessively long assistant turns revisit the same hypotheses without progress. Supported by external judge evaluations, our *loop-aware trajectory filtering* uses maximum per-turn reasoning length as a proxy for looping and improves performance without requiring correct final answers. We find a practical sweet spot at 2,000 training trajectories across 4B and 35B students, with no consistent gains from further scaling. Our 35B agent achieves 66.4% on BrowseComp, 60.8% on BrowseComp-zh, and a 62.8% macro average across six research benchmarks, outperforming the same backbone fine-tuned on the substantially larger S1-DR and Red-Searcher datasets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.