acceptodds
Under review as a conference paper at ICLR 2027

Critic-Grounded Online Fine-Tuning with Active Preference Elicitation

Abstract

The offline-to-online paradigm in reinforcement learning deploys a pretrained actor-critic checkpoint without access to the offline dataset or the reward function that produced it. The data may be withheld for privacy or proprietary reasons, and the reward may have been handcrafted by experts. We study online fine-tuning in this regime, where the only available supervision is preference feedback: a labeler compares two short behavior segments, may abstain, and never provides a per-step reward. Two obstacles arise in this setting. First, a reward learned from preferences is identified only up to a scale and an offset, which the standard Bradley-Terry likelihood cannot pin down, whereas the pretrained critic holds values in the units of the offline reward. Writing an arbitrarily placed reward into the replay buffer corrupts those values and destabilizes fine-tuning. We design a critic gauge, which recovers the reward implied by the frozen critic and regresses the learned reward onto it to fix the scale and offset. Second, preference labels are costly, and the labeler may abstain, so comparisons must be chosen to be both informative and answerable. We score each candidate by the expected reduction in posterior variance across the pool, under a Gaussian posterior over the reward model's parameters, weighted by a calibrated online estimate of the labeler's likelihood of answering. Our selection is sequential, folding each label into the posterior before the next pair is chosen so that redundant comparisons are never selected. On eight standard RL environments, our method closely matches WSRL, a baseline that shares our no-data-retention setting but observes the true online reward, and substantially outperforms PEBBLE, which learns from preferences like ours but starts without a pretrained checkpoint.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.