Actpin: Closed-Loop Semantic-Spatial Control of Generated Motion
Abstract
Sparse spatial control of text-to-motion looks solved: drop an anchor, name the verb, and a guided or inpainted clip meets the point. We measure what that success is made of. When anchors are mined from clips that already perform the verb, a guided controller reads better than text alone (.62 to .92), since the mined anchor narrows generation toward its source performance; yet the pin is met by a snap, the responsible joint accelerating at 148x the clip median, invisible to retrieval. On freely authored anchors the verb collapses (.24-.51 across four families): every optimization these methods run serves space alone, so the point is bought by breaking the verb that no term defends. Actpin closes the loop: on a frozen motion-primitive stream whose motion seeds are the one differentiable input, it seeks per event the seed with legibility an objective measured on the rendered frames and the anchor and travel site both -constraints; across six coupling laws the strict read splits by whether the price remembers (we deploy the closed-form dynamic barrier). The deployed program holds every anchor at 0.7cm and arrival at 0.5cm, reads R@3 .90 / R@1 .71 averaged over three stream draws where the strongest baseline reads .51 / .28, and transfers to a held-out split.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.