Count to Five! Efficient Test-Time Personalization for Pose-Guided Human Video Synthesis
Abstract
Pose-driven human video generation aims to synthesize realistic videos of a person following a specified pose sequence. Existing methods rely on static models that generalize poorly to unseen subjects, leading to appearance drift and mismatched details. We study test-time personalization for pose-guided video diffusion, where a few posed snapshots of the target subject are available at inference. Naive fine-tuning on these snapshots is slow and prone to overfitting, and initial updates can degrade the pretrained model. We therefore propose AdaptAnyone, to our knowledge the first meta-learned framework for test-time personalization of pose-guided video diffusion models. Our instance-level meta-learning strategy treats each training subject as a separate task and learns a parameter-efficient initialization for personalizing a new subject in a few gradient steps. Guided by a sensitivity analysis, we efficiently adapt the components that most influence subject appearance and temporal consistency. With three gradient steps, AdaptAnyone outperforms both the base model and naive fine-tuning, which requires orders of magnitude more steps, in subject fidelity, image quality, and temporal stability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.