Guess What You Ask: Anticipating Moment-level Real Viewer Information Needs in Videos
Abstract
A proactive video agent must decide what information to prepare before a viewer asks. We formulate a prerequisite task of proactive video agent as Moment-level Information Need Anticipation (MINA): given the video context observable up to a video timestamp, predict the questions real viewers are likely to ask. We construct MINA-Data from 12,510 timestamped questions naturally posted by viewers from 872 videos, and evaluate whether an answer to a predicted question would satisfy the observed need. Direct Forecasting is difficult (15.53% at best on Hit@5), even though most generated questions are semantically plausible. This reveals that the challenge is to select which of many plausible needs viewers actually express. We therefore propose MINA-GABT, which predicts a gap-formation state comprising an observed cue, the process by which it evokes an information gap, and the gap itself before verbalizing that gap as a question. A taxonomy induced from training-set viewer behavior guides state prediction. MINA-GABT reaches 21.20% Hit@5 and 27.12% Hit@10, with gains also on a held-out content category. Our analyses show that a video moment can support many plausible questions, but predicting which needs viewers will actually express remains difficult; event-specific gap formation and information beyond the video contribute to this uncertainty.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.