PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
Abstract
Inference-time alignment enables large language models (LLMs) to generate outputs aligned with end-user preferences while keeping the base model fixed. Recent methods use small guidance models to modify token generation under a KL-regularized reward objective. When supervision consists of pairwise preferences, a common approach first estimates a reward model and then learns a decoding guide from its outputs. We introduce PITA, a framework that learns the guidance model directly from preference feedback. PITA uses predicted preference probabilities to reweight the base model's next-token distribution, with iterative data collection and refinement of the guide. We connect preference probabilities to the optimal tilted output distribution and establish performance bounds under linear function approximation. Our experiments cover mathematical reasoning, graph reasoning, and alignment from noisy human preferences. On reasoning tasks, PITA attains accuracy comparable to inference-time alignment methods with access to the true reward function, demonstrating the value of learning guidance directly from preferences.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.