Batch-Coupled Tilted Denoising for Training-free Reward Alignment
Abstract
Inference-time reward alignment steers a pretrained generative model toward a reward without retraining, but gradient-free methods often require costly conditional endpoint sampling to estimate reward-aware corrections. We exploit the fact that generation often occurs in shared-condition batches and introduce cross-trajectory reweighting, which corrects the distribution mismatch so that a con- ditional endpoint sample generated for one trajectory can contribute to the correction of another. This converts endpoint sampling from a per-trajectory cost into a shared batch-level budget under batch coupling, allowing a smaller set of endpoint samples to be reused across trajectories instead of generating them independently. Across multiple backbones and reward objectives, the coupled estimator improves the reward–compute frontier over existing alternatives at matched cost, with the largest gains in low-compute regimes. We further analyze when cross-trajectory reuse is beneficial and the statistical dependence introduced by the coupling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.