acceptodds
Under review as a conference paper at ICLR 2027

Timing Is Not Enough: Mitigating Concurrent Action Omission in Video Generation

Abstract

A prompt for video generation describes actions over time, and a subject is often asked to perform several of them at once. Temporal binding methods decide when each event appears and have made sequential multi-event generation more reliable. Once events contain concurrent actions, however, many requested actions still go missing, and those that appear are often not sustained. We trace the missing actions to two places along the conditioning pathway. In the text encoder, the subject that every frame must read mixes all its events, and isolating events from each other removes the actions that the subject carries. In video-text attention, the actions requested of one subject at the same time compete for the attention at its location, and this attention shifts from one action to another across frames. We propose CoRAct, a training-free approach that keeps the original prompt intact. It anchors each subject to the actions of its current event in the text representation and supports its co-required actions together at every frame in video-text attention. We further construct a benchmark of 84 prompts that specifies an interval for every action. Our protocol judges who performs each action and for how long, and it ranks the compared methods in the same order as human annotators. Across three backbones in image-to-video and text-to-video generation, CoRAct keeps the overall omission rate on par with existing training-free methods while significantly reducing the actions missing from their interval.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.