acceptodds
Under review as a conference paper at ICLR 2027

Designing Coupled Samplers for Language Models

Abstract

Scaling inference compute and reinforcement learning rely on generating batches of language model rollouts. These are typically sampled independently, but prior work on arithmetic sampling shows that this independence can be relaxed: dependent rollouts can be generated in parallel while preserving the marginal distribution. However, this leaves open a consequential design choice—how should the rollouts be coupled? Infinitely many joint distributions preserve the marginals, and classical Monte Carlo suggests this choice matters significantly. In this paper, we establish principles for designing dependent samplers for test-time scaling and reinforcement learning. We first theoretically analyze how the choice of dependent sampler impacts test-time scaling properties. We then characterize these tradeoffs empirically across four controlled reasoning benchmarks. We find that lattice sampling yields the highest coverage, matching i.i.d. pass@ with 25–47% fewer rollouts. In reinforcement learning, stratified sampling performs the best, matching i.i.d performance with 34.5%-59.2% fewer training steps. Our results establish dependent sampling as another design axis for language model inference and learning, and take a step towards a systematic characterization of this design space.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.