acceptodds
Under review as a conference paper at ICLR 2027

Batch-Coupled Tilted Denoising for Training-free Reward Alignment

Abstract

Inference-time reward alignment steers a pretrained generative model toward a reward without retraining, but gradient-free methods often require costly conditional endpoint sampling to estimate reward-aware corrections. We exploit the fact that generation often occurs in shared-condition batches and introduce cross-trajectory reweighting, which corrects the distribution mismatch so that a con- ditional endpoint sample generated for one trajectory can contribute to the correction of another. This converts endpoint sampling from a per-trajectory cost into a shared batch-level budget under batch coupling, allowing a smaller set of endpoint samples to be reused across trajectories instead of generating them independently. Across multiple backbones and reward objectives, the coupled estimator improves the reward–compute frontier over existing alternatives at matched cost, with the largest gains in low-compute regimes. We further analyze when cross-trajectory reuse is beneficial and the statistical dependence introduced by the coupling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.