OVAL: Online Value Adaptation from Lookahead for Test-Time Alignment
Abstract
Test-time alignment (TTA) increasingly relies on inference-time search to steer frozen language models toward desired preferences. Yet existing planning methods under-utilize the information generated by search. Lookahead rollouts expose how candidate segments differ in their downstream preference outcomes, but these comparisons are typically consumed only by the current decision and then discarded, leaving subsequent decisions to the same critic. We introduce OVAL, a novel TTA framework that turns lookahead rollouts from a transient search signal into a source of online value adaptation. For each prompt, OVAL propagates rollout-induced rankings back to candidate segments and accumulates them as prompt-specific supervision. It then employs Follow-the-Regularized-Leader (FTRL) to learn a lightweight residual correction to an offline critic, enabling the critic to progressively incorporate evidence collected along the generation trajectory. Theoretically, under informative prompt-level signals, cross-step evidence accumulation changes the decision-regret regime from linear to bounded, while OVAL’s FTRL update achieves predictive regret. Experiments across multiple LLM backbones, benchmarks, and preferences demonstrate consistent gains over strong alignment baselines, with ablations confirming the importance of prompt-specific critic adaptation and cross-step evidence retention. This work highlights online value adaptation as a general principle for TTA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.