Scope2Rubric: Recovering Query-Specific Evaluation Scope before Rubric Generation for Long-Form Deep Research
Abstract
Long-form deep research (LFDR) tasks lack definitive answers for RL with verifiable rewards (RLVR), motivating rubric-based rewards for training deep research agents. Rubric effectiveness hinges on jointly covering full task requirements and distinguishing task-relevant quality differences among reports. Yet existing methods often miss implicit requirements or capture task-irrelevant response differences. We attribute both failures to an overlooked constraint: full task requirements are not explicitly elicited to constrain what rubrics should cover and which response differences they should capture. Accordingly, given the multifaceted nature of LFDR tasks, our key insight is to group requirements sharing an evaluation goal into directions that jointly form a task-defined evaluation scope, constraining both rubric generation and decomposition. We propose Scope2Rubric (S2R), a two-stage rubric construction framework. Stage I extracts full task requirements from query and reference analyses, organizes them into weighted directions, and reviews their coverage to form the evaluation scope. The scope constrains rubric generation within each direction, followed by individual, intra-direction, and cross-direction refinement to ensure the rubric set’s necessity, completeness, and non-redundancy. Stage II identifies rubrics that fail to distinguish candidate responses and decomposes them by contrasting task-relevant differences within the fixed Stage I scope, preventing salient differences from inducing task-irrelevant rubrics. Extensive experiments show S2R achieves the best human-preference alignment over strong baselines. When used as rewards to train Qwen3-8B, it yields a superior 67.72 average score across five LFDR benchmarks. Further analyses confirm these gains stem from structured task-defined scope constraints.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.