acceptodds
Under review as a conference paper at ICLR 2027

Fast & efficient online alignment with multidimensional reward model

Abstract

We consider online alignment of language model outputs to user-specific preferences. We assume each user's preferences are driven by an unknown function over a chosen multi-dimensional vector of measurable text qualities (e.g., helpfulness, conciseness, factuality), so that personalisation reduces to learning that function from live feedback. Online adaptation without updating the underlying LLM requires learning this function from sparse and heterogeneous feedback, with no guarantee that any structural assumption about how qualities combine fits every user. Under a KL-regularised objective, a standard identity gives the optimal policy as an exponentially tilted base distribution whose tilt is the user's reward; with it assumed to be a function of qualities, learning the user's preferences therefore reduces to learning that reward from feedback. We cast this as a bandit problem and solve it with Thompson sampling under two posterior families: a linear- Gaussian belief and a non-parametric Gaussian process (GP). In synthetic experiments, the linear variant converges within tens of interactions where its assumption holds, and the GP variant succeeds on non-linear structures where the linear model degrades. On both the synthetic worlds and a PRISM-derived persona benchmark, achieved reward depends jointly on what the posterior has learned and on how sharply the decoder acts on it. We characterise this dependence and show its role alongside the choice of quality basis in the gap to an in-context baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.