Editable Action Words for Generative Motion Editing
Abstract
Text-based generative motion editing revises a generated clip from language, but mostly within an action, changing how it is performed. Editing the set of actions itself (adding, switching, moving, or removing one) has no direct handle: a generator returns a uniform grid of frames or tokens in which no action exists as an element, so the edit is approximated—a torn boundary, a disturbed rest, a sliding foot. We instead generate the units editing needs: a motion is a set of action words, one per action per body-part track. Each is a Gaussian on the clip's clock (a learned center, duration, and opacity) carrying a text-aligned content vector (what) and a manner residual (how); a named action is the group of action words that share it. The action is now a first-class element: the caption reads back as a legible timeline and an edit is a set operation that re-renders in distribution. Across insert, switch, move, and delete on an action-level editing benchmark, action words alone realize the operation while preserving the untouched motion and keeping foot contact physically plausible, where masked-infilling, attention, and instruction-tuned editors each give up one of these.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.