acceptodds
Under review as a conference paper at ICLR 2027

Direct Trajectory Optimization for Inference-Time Alignment of Text-to-Image Diffusion Models

Abstract

Alignment of text-to-image diffusion models with certain reward functions is a key challenge in controllable generation, yet existing approaches often rely on optimizing model inputs or fine-tuning denoising networks, which requires backpropagation through large denoising models and can be computationally expensive at inference time. In this work, we propose Direct Trajectory Optimization (DTO), a training-free framework that performs alignment by directly optimizing the sampling trajectory of diffusion models, treating intermediate denoising scores and sampled noise as optimization target variables without requiring gradients through the denoising network. To mitigate out-of-distribution problem, we introduce Reward-Optimized Score Guidance, which interpolates between optimized scores and pretrained model predictions to stabilize sampling while theoretically provably preserving reward improvement generally. Empirically, DTO improves alignment performance across three diffusion models on two benchmarks, outperforming existing inference-time alignment methods while maintaining high efficiency. The results suggest that trajectory-level optimization is a promising direction for efficient and effective diffusion model alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.