acceptodds
Under review as a conference paper at ICLR 2027

One Video, Many Viewers: A Benchmark for Persona-Conditioned Video Highlights and Chapter Titles

Abstract

Highlights and chapters help viewers navigate long videos, yet they are usually produced once and shared by all viewers. Which moments matter and how each part should be described can differ between viewers, but existing evaluations rarely test whether these outputs fit a particular viewer. We study highlight selection and chapter titling conditioned on a viewer persona, a short description of a viewer's background and interests. Evaluating this setting requires annotations for many personas on the same video. Large language models make such annotations easy to produce, but differences between annotations may reflect run-to-run variation rather than the persona. We introduce One Video, Many Viewers, a benchmark with synthetic annotations for 259 source videos across highlight selection and chapter titling, built with 24 census-grounded synthetic personas. Highlight annotations are guided by a video-specific perspective derived from the persona, and chapter titles are generated for fixed intervals. Repeated-generation diagnostics show that persona-associated differences recur within an annotation model, whereas persona-specific title contrasts are much smaller across models than within each model. Evaluating existing models without benchmark-specific training, we find small but consistent gains over another persona in base AP for highlight models that score every unit. In contrast, we find no consistent gains for models that select units or for chapter titling. These results show that agreement with persona-conditioned references should be read alongside comparisons with other-persona and persona-free conditions. We release the annotations, personas, and perspectives for viewer-specific evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.