acceptodds
Under review as a conference paper at ICLR 2027

Internalizing Rubric Guidance for Research Agents

Abstract

A research rubric specifies what an answer should cover, but using it only as a scalar reward leaves its instructional content unused. We study how to turn this content into better research agents that act without rubric hints. We introduce rubric-guided self-distillation (RSD): criterion-level feedback selects weak queries, the current policy generates new tool trajectories with targeted rubric guidance, and a reward threshold filters the resulting demonstrations. The agent then learns these trajectories under the original, unhinted prompt in supervised phases interleaved with fresh-policy reinforcement learning. No separate demonstration model is required. On Qwen3-8B, RSD improves over rubric-based GRPO by 14.23, 7.61, and 2.84 score points on ResearchQA, DeepResearch Bench, and DeepResearch Bench II, with paired 95% task-bootstrap confidence intervals excluding zero on all three benchmarks. The method also improves LSB from 26.04 to 28.79. At 14B and 32B, it outperforms both starting models and size-matched GRPO on all four research benchmarks, with RQA gains over GRPO of 15.07 and 14.47 points. Strategy comparisons highlight rubric-guided demonstration learning as a major contributor to the gains, and RSD achieves the highest scores on all four research benchmarks among the evaluated 8B configurations. These results show that rubrics can provide not only rewards but also demonstrations that improve research performance without rubric hints at inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.