acceptodds
Under review as a conference paper at ICLR 2027

Structure-Anchored Conditional Response Steering for Human Video Generation

Abstract

Structurally conditioned human video generators already follow trajectories strongly constrained by pose, identity, and reference appearance, making unrestricted classifier-free guidance (CFG) prone to perturbing otherwise useful structural information. We revisit guidance from a conditional-response perspective and distinguish target-specific semantic and structural responses through selective changes of the conditioning configuration while keeping the remaining context fixed. Based on these axis-wise responses, we propose a structure-anchored steering formulation in which the model's own local response geometry determines where intervention is applied, while a bounded intervention budget regulates semantic steering in direction and diffusion time before the structural response is recomposed with controlled output magnitude. On Wan-Animate, our method achieves the best numerical results among the evaluated guidance methods on CLIP, T2V, MPJPE, aggregate VBench2 Human Fidelity Score, and both VideoScore2 visual-quality and physical-consistency criteria. Cross-backbone evaluation on other DiT-based algorithms further examines the transfer of the proposed guidance principle. These results suggest that guidance for structured human video generation is better formulated as controlled conditional-response steering than as unrestricted transformation of a single conditional discrepancy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.