acceptodds
Under review as a conference paper at ICLR 2027

LatticeSMC: Where to Spend Inference-Time Compute in Chunked Sequence Generators

Abstract

Long-form generators for music, motion and video produce a sequence chunk by chunk, each chunk by iterative denoising, while the rewards that matter are defined on the whole sequence. Existing inference-time steering methods act on one axis of this process at a time, best-of-N at the end, Feynman-Kac steering across denoising steps, streaming pipelines across chunks, and have been compared under different return rules at unmatched compute. We introduce budget-matched chunked steering, the task of maximizing a sequence-level reward with a chunked generator under a fixed number of denoiser evaluations, and propose LatticeSMC, a sampler derived from a Feynman-Kac model on the two-dimensional lattice of chunk index and denoising step. Two exact telescoping results make the design derivable: for chunk-additive rewards the two axes carry identical weights, so resampling belongs wherever lookahead is cheapest; for terminal rewards any prefix score yields an exact intermediate potential, so a prefix-evaluable reward is a twist with no estimation and no extra denoiser evaluations. LatticeSMC resamples properly on these exact potentials, at chunk boundaries and, when the plug-in reward is free, within chunks, and returns either a draw from its particle approximation of the reward-tilted target or its best particle, and recovers Feynman-Kac steering as its within-chunk schedule. Under matched compute, on a music-to-dance diffusion model and a 40-second text-to-music model, it raises beat alignment from 0.234 to 0.441 (best-of-N: 0.354) and prompt adherence from 0.470 to 0.560 at 32 particles with held-out quality metrics at base, holds its lead on the long-range reward at six and eight chunks, and is preferred by human raters over every comparator in 60 to 77 percent of pairwise judgements. Two measurable properties of the reward settle what remains: how hard to commit at a boundary follows the information in its exact potential, which the effective-sample-size gate reads off and greedy pruning cannot, and whether content-axis lookahead pays follows the within-set predictability of future reward, measured before any particle is spent.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.