acceptodds
Under review as a conference paper at ICLR 2027

Do Not Stop at Frame Scoring: Learning to Compose Video Summaries

Abstract

Video summarization is a task to extract a sequence of segments to preserve the key content of a video. Recent work on this task tends to focus only on predicting frame importance scores, leaving the segmentation and budget-constrained selection as a fixed pipeline. However, our empirical analysis reveals that a more precise importance score does not reliably translate into a better summary, suggesting that summary composition itself should be trained. Directly learning to compose a summary, however, requires structured summary-level supervision, which is prohibitive to obtain at scale. To address this bottleneck, we introduce offline pseudo-summary construction, transforming existing frame-importance annotations and video content into a structured, budget-aware pseudo-summary. We then introduce SumComposer, a model to autoregressively add a segment to the summary, conditioned on preceding selections. Once trained, it requires neither importance annotations nor language model calls to summarize a video. Our experiments verify that our pseudo-summaries are better aligned with human annotations than fixed construction, and SumComposer trained on these pseudo-labels substantially improves matching and boundary alignment over fixed composition.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.