Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering
Abstract
Linking people’s appearance and actions to their character identities is essential for understanding video narratives. We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation of video-language models. Starting from movie clips in LSMDC v2, our pipeline matches detected faces to actor reference images, tracks characters across frames, and constructs shot-level inputs with identity-linked bounding boxes. We use a strong vision-language model to generate identity-aware captions and associated questions, followed by manual verification and filtering to obtain an evaluation benchmark of 750 captioned clips and 3,000 person-centric questions. We investigate five grounding strategies combining textual coordinates with visual face or estimated person boxes, and evaluate them across multiple Video-MLLM families at approximately 2B, 4B, and 8B parameters, as well as larger frontier models. Combining visual face boxes with textual coordinates provides the most consistent performance across scales and significantly improves question-answering accuracy over coordinates alone. We further observe that smaller models tend to over-assign known identities when the queried person is not grounded, whereas larger models better distinguish such UNIDENTIFIED cases. Building on the strongest grounding formulation, we introduce BAC by LoRA fine-tuning Qwen models at 2B, 4B, and 8B scales on approximately 32K identity-aware captioned clips. BAC-8B reaches 93.20% overall QA accuracy, ranking behind only GPT-5.6 Sol among the evaluated frontier models while outperforming the other open-weight and proprietary systems. Overall, our results show that explicitly communicating who is where, together with lightweight task-specific adaptation, substantially improves identity-aware video understanding without modifying the underlying model architecture. We publicly release the training data, evaluation benchmark, source code, and trained BAC checkpoints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.