HarnessRL: Co-Evolving Policies with a Just-in-Time Training Harness for Open-Ended RL
Abstract
Rubric-based reinforcement learning enables open-ended generation using natural-language criteria as rewards, but these signals can become ineffective as the policy evolves. We identify two failure modes: non-discriminative rewards, where scores collapse to uniformly low or high values due to insufficient exploration or a mismatch between task difficulty and policy capability, and spurious rewards, where higher scores do not correspond to better responses. Both reflect a fixed training configuration that fails to adapt as the policy evolves, including what to train on, how to explore, and how to evaluate. We introduce HarnessRL, a training-time harness that co-evolves with the policy through three interfaces: task sampling, rollout guidance, and rubric criteria. HarnessRL maintains a training-time memory that retrieves within-task history and cross-task evidence. A harness evolver uses this evidence to attribute each reward failure to the interface whose update can address it, then runs a proposer and just-in-time critic loop in which the critic accepts a candidate update only after measuring its effect on the current frozen policy. We further introduce Rubric-Mix-10k, a unified five-domain training and evaluation suite for open-ended RL spanning medical, science, writing, role-playing, and instruction-following tasks. On Rubric-Mix-10k, HarnessRL outperforms the strongest baselines by 2.8 to 3.1 points across both 4B and 9B model scales and nearly halves the training steps to its best checkpoint.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.