ProxyUp: Training-Free Generation of Coherent Video Interactions from Motion Proxies
Abstract
Motion proxies can specify an action without depicting the objects, contacts, and scene needed to realize it naturally. We study how to turn such incomplete guidance into visually natural videos with coherent interactions while retaining the intended actions and key event structure. We propose ProxyUp, a training-free framework built on a pretrained video generator. Its inference design combines region-wise proxy initialization with full-latent refinement using predict-and-perturb updates, followed by ODE sampling. The mask selects initial guidance but does not restrict subsequent updates: the guided object and new content can adapt together at contacts and occlusions. The proxy thus guides generation without becoming a framewise target. On simulation and real-world proxies, ProxyUp achieves higher event-based motion and mechanics scores, imaging quality, and mean human ratings than the evaluated baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.