acceptodds
Under review as a conference paper at ICLR 2027

Sound Sparks Motion: Audio and Text Tuning for Video Motion Editing

Abstract

Motion-centric video editing remains difficult for large generative video models, which often respond well to appearance changes but struggle to produce specific, localized actions or state transitions in an existing clip. We introduce Sound Sparks Motion, a training-free framework that enables motion editing in an audio-visual video generation model by tuning its internal multimodal conditioning signals at test time. Rather than modifying model weights, our method tunes two lightweight variables: the text conditioning and the audio latent. We find that tuning the text alone is often insufficient, as audio and motion are highly entangled in the underlying model. By also optimizing the audio latent, we can more effectively steer the desired motion, while the resulting audio can simply be discarded after generation. Since there is no direct way to evaluate temporal alignment between text and motion, we guide the tuning process using a vision–language model that provides feedback indicating whether the intended motion appears in the generated video. This simple supervision yields an effective semantic objective for motion editing, while regularization and perceptual-temporal constraints help preserve content and visual quality. Beyond per-video tuning, we show that the learned latent controls are transferable across videos, suggesting that they capture reusable motion-edit directions rather than overfitting to a single example. Our results highlight multimodal conditioning tuning, particularly through the audio pathway, as a promising direction for motion-aware video editing, and suggest that test-time tuning can serve as a lightweight probing mechanism that helps reveal latent motion controls embedded in the model’s multimodal conditioning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.