acceptodds
Under review as a conference paper at ICLR 2027

HarnessRL: Co-Evolving Policies with a Just-in-Time Training Harness for Open-Ended RL

Abstract

Rubric-based reinforcement learning enables open-ended generation using natural-language criteria as rewards, but these signals can become ineffective as the policy evolves. We identify two failure modes: non-discriminative rewards, where scores collapse to uniformly low or high values due to insufficient exploration or a mismatch between task difficulty and policy capability, and spurious rewards, where higher scores do not correspond to better responses. Both reflect a fixed training configuration that fails to adapt as the policy evolves, including what to train on, how to explore, and how to evaluate. We introduce HarnessRL, a training-time harness that co-evolves with the policy through three interfaces: task sampling, rollout guidance, and rubric criteria. HarnessRL maintains a training-time memory that retrieves within-task history and cross-task evidence. A harness evolver uses this evidence to attribute each reward failure to the interface whose update can address it, then runs a proposer and just-in-time critic loop in which the critic accepts a candidate update only after measuring its effect on the current frozen policy. We further introduce Rubric-Mix-10k, a unified five-domain training and evaluation suite for open-ended RL spanning medical, science, writing, role-playing, and instruction-following tasks. On Rubric-Mix-10k, HarnessRL outperforms the strongest baselines by 2.8 to 3.1 points across both 4B and 9B model scales and nearly halves the training steps to its best checkpoint.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.