Do Not Stop at Frame Scoring: Learning to Compose Video Summaries
Abstract
Video summarization is a task to extract a sequence of segments to preserve the key content of a video. Recent work on this task tends to focus only on predicting frame importance scores, leaving the segmentation and budget-constrained selection as a fixed pipeline. However, our empirical analysis reveals that a more precise importance score does not reliably translate into a better summary, suggesting that summary composition itself should be trained. Directly learning to compose a summary, however, requires structured summary-level supervision, which is prohibitive to obtain at scale. To address this bottleneck, we introduce offline pseudo-summary construction, transforming existing frame-importance annotations and video content into a structured, budget-aware pseudo-summary. We then introduce SumComposer, a model to autoregressively add a segment to the summary, conditioned on preceding selections. Once trained, it requires neither importance annotations nor language model calls to summarize a video. Our experiments verify that our pseudo-summaries are better aligned with human annotations than fixed construction, and SumComposer trained on these pseudo-labels substantially improves matching and boundary alignment over fixed composition.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.