acceptodds
Under review as a conference paper at ICLR 2027

ARRCE: Learning Process Rubrics from Static Anchors for Open-Domain Reinforcement Learning

Abstract

We introduce Asymmetrically Anchored Rubric–Response Co-Evolution ARRCE, a reinforcement learning framework that learns process rubrics from automatically synthesized static supervision for open-domain tasks. The central difficulty is to adapt evaluation criteria to the policy's changing behavior while preserving the task requirements that define a useful answer. ARRCE decouples rubric synthesis into an offline stage that constructs and freezes task and evidence-backed fact criteria, and an online stage that proposes process criteria from sampled trajectories whose relative strengths are determined by static-rubric scores. The static criteria also supervise candidate quality: a proposed process rubric is rewarded for separating held-out trajectories in a direction consistent with static performance, subject to observability and scope checks, before entering a query-specific memory for subsequent response rewards. A shared policy learns response generation and rubric proposal through separate objectives in one parameter update, yielding a synthesis loop that requires no additional human annotation of trajectory preferences or process rubrics. We instantiate this framework on dataset-derived tasks with synthesized static rubrics across healthcare, scientific question answering, and long-form writing. The experimental results show that full ARRCE improves over the response-SFT baseline by on ResearchQA (), OP points on LLMEval-Med (), on WritingBench (), and on FollowBench (), and the gains are positive across the four benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.