The Router Is Already in the ARMs: Disagreement-Guided Multi-Objective Test-Time Alignment for LLMs
Abstract
Multi-objective test-time alignment uses autoregressive reward models (ARMs) to steer a frozen language model toward user-specified preferences. But when should reward guidance matter, and how strongly should it override the base policy? We show that the answer is already in the ARMs: disagreement among objective-conditioned next-token distributions reveals cross-reward forks, where different local reward directions can lead to meaningfully different downstream reward outcomes. This disagreement has a precise information-theoretic interpretation—generalized Jensen–Shannon divergence equals the mutual information between objective identity and the next token—and is strongly correlated with downstream cross-reward separation (Spearman 0.766). Building on this insight, we introduce Disagreement-Informed Routing (DIR), which uses ARM disagreement to modulate the token-wise step size from the base policy along the requested preference direction. DIR requires no learned router, auxiliary model, or policy fine-tuning. Experiments across safety alignment, cross-domain transfer, and model scales demonstrate effective multi-objective alignment using the routing signal already present in the ARMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.