acceptodds
Under review as a conference paper at ICLR 2027

Steer Early, Compose Better: Avoiding Late Horizon Steering by Optimal Residual Control for Diffusion Alignment

Abstract

Existing diffusion-alignment methods struggle with compositional tasks such as object-count constraints, object-attribute matching, and concept unlearning. We conjecture that this limitation arises from inaccurate estimates of intermediate-state values by the auxiliary networks that propagate rewards from clean samples to intermediate states. We demonstrate that such parameterizations exhibit Late Horizon Steering, a phenomenon in which steering is observed only near the clean endpoint, limiting their effectiveness on tasks that require intervention during the early stages of generation. To mitigate this problem, we build on a stochastic optimal-control formulation by exploiting the relation between the optimal control and the score residual: the difference between the controlled score field and the pretrained score field. Using the Hamilton-Jacobi-Bellman dynamics of the value function, we derive a coupled SDE governing the evolution of the diffusion state and its corresponding optimal score residual during the generation process, which we call the state-residual SDE. Based on this characterization, we propose Residual Consistency, a learning objective that enforces consistency with a discretization of the state-residual SDE, enabling alignment without an auxiliary network. We demonstrate that our proposed method achieves a superior trade-off across broad classes of post-training objectives, including human-preference, compositional, and safety alignment, achieving up to 12% improvement over baselines on compositional tasks and up to 20% lower detection rate on unlearnt concepts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.