RubricAtlas: Decoupling Rubric Discovery from Fact Validation to Optimize Generalist Deep Research Agents
Abstract
Deep research agents are increasingly deployed to solve complex tasks by dynamically searching the web, integrating multi-source evidence, and composing long-form reports. Since such reports lack unique reference answers, task-specific rubrics are crucial for steering agent capability enhancement. However, existing rubric synthesis paradigms—such as offline model comparisons or online rollout evaluations—accept newly generated criteria without independent verification. This induces severe reward hacking, encouraging policies to favor excessive verbosity, superficial reasoning, and factually ungrounded demands, while existing datasets and queries suffer from limited domain coverage. To overcome these limitations, we present **RubricAtlas**, a unified framework for optimizing deep research agents across diverse knowledge domains. RubricAtlas builds a four-level taxonomy with 234 leaf concepts for diverse task synthesis. To enforce reliable supervision, we introduce Decoupled Rubric Discovery and Validation (DRDV), which decouples depth-oriented criteria discovery from search-augmented factual verification to filter ungrounded constraints. A Memo Context Manager maintains structured findings and evidence provenance throughout training and inference, supporting context integrity over long research horizons. For policy optimization, Group reward-Decoupled Normalization Policy Optimization (GDPO) normalizes rubric-wise advantages before aggregation, preserving fine-grained reward distinctions across varying reward scales. Extensive experiments show that RubricAtlas, with Qwen3-14B as its backbone, matches or surpasses strong proprietary models across three evaluated deep-research benchmarks, achieving an **82.60%** win rate on DeepConsult, a **46.67** RACE score on DeepResearch Bench, and **62.15%** rubric compliance on ResearchRubrics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.