acceptodds
Under review as a conference paper at ICLR 2027

LET RHYTHM TELL YOUR STORY

Abstract

In audio-driven video editing, jointly determining ordered video clips' relative durations and positions on a given music timeline is a cross-modal structured prediction problem with multiple valid solutions and global constraints. We formulate this as constrained boundary-trajectory prediction, jointly specifying relative clip durations and timeline positions. Our primary evaluation uses a given start time. We introduce Cutalio, a unified framework for trajectory data construction, modeling, inference, and relational evaluation. Using real edits and independent-start augmentations, we train a multimodal masked model that fuses video, music, and rhythmic evidence. The model jointly predicts unknown boundaries; differentiable dynamic programming computes distributions over feasible paths under configurable outer constraints, including per-clip available source durations, an optional target edit duration, and known cuts, and feeds these distributions back into subsequent predictions. At inference, the framework can generate a full timeline from the given start or complete it around fixed cuts. Progressive reveal, multi-trajectory sampling, and model-based ranking provide test-time scaling. Because absolute cut error against a single reference cannot adequately assess multiple valid arrangements, we introduce pairwise relational metrics for duration ordering, duration ratios, local acoustic changes, and sequence-level acoustic trends. Experiments show that music content and rhythmic cues are complementary, known anchors aid completion, and independent-start augmentation helps primarily under multi-trajectory inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.