OACap: Observation-Adaptive Video Captioning via Caption-Centered Agentic Evidence Grounding
Abstract
Dense video captioning offers a rich textual interface to video understanding and generation. Recent advances in multimodal large language models have substantially improved fine-grained video perception and long-form text generation, laying a strong foundation for dense video captioning. However, existing methods still face two major challenges: achieving comprehensive coverage of fine-grained video content, and maintaining factual grounding and global semantic coherence as captions grow longer and more detailed. To this end, we propose Observation-Adaptive Video Captioner (OACap), an agentic model that learns to think with video. Rather than relying on a fixed perceptual context, OACap treats the evolving caption itself as the state of reasoning: it identifies missing, unsupported, or semantically inconsistent content, and adaptively acquires additional visual, temporal, or external evidence to verify and refine the description.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.