ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents
Abstract
Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad. Rubric-based rewards improve interpretability through natural-language criteria, but many existing approaches score complete trajectories with a frozen closed-source judge, limiting step-level guidance and adaptation during training. We propose ARCO (Adaptive Rubric CO-evolution), which generates interpretable, action-specific per-step rubrics and dense rubric-conditioned rewards, then continuously co-evolves the rubric model with the agent on shared on-policy rollouts. After warmup, trajectory decomposition fits step rewards to terminal outcomes without step labels or an external judge during online training. Across HotpotQA, 2WikiMultiHopQA, and MuSiQue with two open-source backbones, ARCO achieves the highest EM in all settings over outcome-, rubric-, and process-reward baselines. We further analyze ARCO along several dimensions—the action-specificity of its rubrics, the number of criteria per step, and the rubric model's scale and family—showing that it yields an interpretable step-level signal useful for diagnosing agent behavior. Code and data are available at https://anonymous.4open.science/r/arco-anon-2C61.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.