acceptodds
Under review as a conference paper at ICLR 2027

PIGEONS: Incremental Long-Video Captioning with Persistent Identity Gallery and Evolving Overarching Narrative Summary

Abstract

Detailed video captioning has become an essential semantic interface for multimodal understanding and generation. While existing methods mainly target short videos, extending them to long videos remains challenging, primarily due to limited context windows of current video understanding models. A naive solution is to adopt segment-wise captioning followed by concatenation, yet this paradigm often suffers from inconsistent subject references and fragmented long-range narratives. To mitigate these issues, we introduce PIGEONS, an incremental long-video captioning model built upon a Persistent Identity Gallery and an Evolving Overarching Narrative Summary. Before caption generation, PIGEONS scans the entire video to construct an identity gallery that captures key subjects along with their visual references, thereby providing a stable referential anchor. The model then performs clip-by-clip captioning conditioned on the persistent gallery and an evolving overarching narrative summary, which compactly maintains and updates essential long-range contextual information after each captioning step. To enable these capabilities, we develop a three-stage training strategy that integrates supervised fine-tuning and reinforcement learning. Extensive experiments demonstrate that PIGEONS significantly improves long-video captioning performance, achieving superior accuracy, completeness, and subject referential consistency over existing approaches, while maintaining competitive performance on short-video captioning. PIGEONS will be open-sourced to facilitate future research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.