acceptodds
Under review as a conference paper at ICLR 2027

ActDiff: Structured Action Search for Test-Time Alignment of Diffusion Models

Abstract

Diffusion models achieve strong generative performance, yet aligning their outputs with human preferences remains challenging. Test-time scaling provides a flexible alternative to fine-tuning by allocating additional inference compute to reward-guided search without retraining. However, existing approaches primarily explore stochastic noise perturbations, offering a narrow, unstructured search space that becomes increasingly restrictive for distilled few-step and deterministic samplers, where stochastic branching is limited or absent. We examine this limitation from an information-theoretic perspective, showing that reward improvement is constrained by the influence of the searched variable on the final output. Motivated by this analysis, we introduce ActDiff, a plug-and-play framework that searches over structured control conditions, such as prompt conditioning and guidance strength, using lightweight, state-dependent diagnostics from the denoising process to guide exploration. This enables more systematic and reward-relevant trajectory variations than noise-only search. ActDiff is gradient-free and compatible with diverse reward functions, search strategies, and diffusion samplers, covering both stochastic and deterministic sampling with standard and distilled few-step models. Experiments demonstrate that expanding the search space beyond noise consistently improves reward alignment and generation quality over noise-centric baselines, with particularly pronounced gains in distilled few-step settings, where noise-based exploration quickly saturates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.