GraphRover: Structured Post-Training for Scientific Agents
Abstract
To work like scientists, agents must gather useful evidence and provide accurate, complete explanations. Yet imitating successful trajectories can reinforce unnecessary searches, while rewarding final conclusions alone can overlook missing facts and experimental conditions. We introduce GraphRover, a graph-guided post-training framework for scientific agents. First, Graph-Guided Inquiry Distillation constructs an operation–evidence graph to identify how tool interactions inform later decisions and answer formation. It masks imitation loss for confidently redundant or unproductive actions while preserving the full interaction history, allowing agents to learn useful behavior without losing the context behind it. Second, Completeness-Aware Reinforcement Learning uses rubrics grounded in reference answers and retrieved evidence to specify the findings, conditions, and conclusions required by each question. Its reward combines item-level answer quality with the coverage of adequately satisfied requirements, encouraging accurate scientific explanations with fewer omissions. Across five benchmarks spanning scientific reasoning, research, and deep search, GraphRover improves its base model’s average score from 38.71 to 55.96. It outperforms both Qwen3.5-35B-A3B and Qwen3.6-35B-A3B on three benchmarks, exceeding the stronger 35B reference on each by 4.50 percentage points on FrontierScience-Olympiad, 9.10 percentage points on FrontierScience-Research, and 7.00 F1 points on DeepSearchQA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.