acceptodds
Under review as a conference paper at ICLR 2027

Directions That Matter: Efficient Online RLHF via Preference Geometry Sketching

Abstract

Online Reinforcement Learning from Human Feedback (RLHF) enables models to continually collect preference feedback and refine their reward models during interaction. However, scaling existing online RLHF methods to large language models (LLMs) remains computationally challenging. Their one-pass estimators rely on high-dimensional local curvature to preserve statistical efficiency, leading to substantial computation and memory costs that grow at least quadratically with the feature dimension. We address this challenge by proposing a geometry-sketched reward modeling framework that reduces the dependence on the feature dimension to linear. The key idea is to approximate the geometry of accumulated preference data using matrix sketching: we retain the dominant directions of preference geometry without explicitly maintaining the full curvature matrices. We establish theoretical efficiency–utility trade-offs for representative online RLHF settings with different data-collection strategies and forms of comparison feedback. Extensive experiments show that our method provides an effective efficiency–utility trade-off, consistently improves over existing approximation baselines, and scales to larger LLMs with limited additional computational overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.