STAR: Selective Timestep Anchor Reward for Identity-Preserving Video Generation
Abstract
Identity-preserving image-to-video post-training must strengthen facial identity without unnecessarily suppressing motion. We introduce STAR, which addresses this trade-off by selectively placing reward gradients along the denoising trajectory. It calibrates an anchor window from the joint response of identity similarity and facial dynamics to timestep-wise perturbations. Within a single-leap backpropagation framework, it samples anchors from this window while an identity-forward connector evaluates the reward on the fully generated video. The resulting update separates timestep coverage from backpropagation depth. Our analysis characterizes terminal scoring, endpoint attenuation of raw reward gradients, and the local identity–motion trade-off of anchor selection. We build HQ4000, a 4,000-clip dataset with per-subject pose and expression annotations, and adapt an identity verifier to support training. We evaluate STAR on four I2V settings: HunyuanVideo-1.5 with 50 and 12 sampling steps, Wan2.2 with 50 steps, and MiniMax-H3. STAR improves test500 identity similarity by 9.8%–35.6% over the base models, with facial dynamics well preserved and no visible reward hacking. Our project page is available at https://anonymous.4open.science/w/STAR-2027/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.