Structure-Anchored Conditional Response Steering for Human Video Generation
Abstract
Structurally conditioned human video generators already follow trajectories strongly constrained by pose, identity, and reference appearance, making unrestricted classifier-free guidance (CFG) prone to perturbing otherwise useful structural information. We revisit guidance from a conditional-response perspective and distinguish target-specific semantic and structural responses through selective changes of the conditioning configuration while keeping the remaining context fixed. Based on these axis-wise responses, we propose a structure-anchored steering formulation in which the model's own local response geometry determines where intervention is applied, while a bounded intervention budget regulates semantic steering in direction and diffusion time before the structural response is recomposed with controlled output magnitude. On Wan-Animate, our method achieves the best numerical results among the evaluated guidance methods on CLIP, T2V, MPJPE, aggregate VBench2 Human Fidelity Score, and both VideoScore2 visual-quality and physical-consistency criteria. Cross-backbone evaluation on other DiT-based algorithms further examines the transfer of the proposed guidance principle. These results suggest that guidance for structured human video generation is better formulated as controlled conditional-response steering than as unrestricted transformation of a single conditional discrepancy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.