Learning Stratified Policies for Inference-Time Objectives
Abstract
Inference-time scaling improves the performance of large language models by sampling multiple responses and evaluating their collective utility. However, many existing approaches obtain these samples by repeatedly sampling from a single policy. In this work, we study if the sample budget at inference time can instead be allocated across a population of policies. We introduce stratified policies, where each of learned policies contributes one response to a -sample inference-time objective. We realize this policy population using separate LoRA adapters on a shared language model backbone and optimize it for pass@. We theoretically motivate the approach by analyzing the linear setting; a population of LoRAs can achieve zero best-of- approximation error in settings where every single LoRA has strictly positive prediction error. Empirically, we demonstrate that stratified policies consistently improve pass@ over a single-policy baseline across multiple reasoning tasks. Overall, our results highlight how distributing samples across complementary policies can be a more effective use of sample budget than repeated i.i.d. sampling from a single policy at inference time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.