acceptodds
Under review as a conference paper at ICLR 2027

Beyond More Samples: Test-Time Scaling of Training-Free Guidance in Generative Models

Abstract

Recent work on test-time scaling has demonstrated that additional computation during inference can improve the quality of generative model outputs. In parallel, training-free guidance has enabled pretrained diffusion and flow-matching models to pursue specific target objectives during sampling without retraining. These two inference-time approaches can complement one another in improving a pretrained sampler's outputs for the target objective. We study their intersection through three recurring scaling strategies applied to guided sampling: (1) averaging local draws, (2) propagating parallel sampling trajectories, and (3) selecting among final samples. Yet alongside inconsistencies reported in past research, we find that these strategies often fail to reliably improve final outputs and can even severely degrade them. The strategy that improves outputs most also changes, even between samplers that share the same checkpoint and objective. In this work, we develop a common characterization of how additional compute acts in each scaling strategy: what increasing the sampling budget actually refines, why that refinement can still fail to improve final outputs, and where to intervene when further scaling stalls. Building on this characterization, we also show that conventional summaries used to monitor these strategies, such as weight concentration or rank agreement, can omit information needed to predict the effects of scaling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.