acceptodds
Under review as a conference paper at ICLR 2027

OACap: Observation-Adaptive Video Captioning via Caption-Centered Agentic Evidence Grounding

Abstract

Dense video captioning offers a rich textual interface to video understanding and generation. Recent advances in multimodal large language models have substantially improved fine-grained video perception and long-form text generation, laying a strong foundation for dense video captioning. However, existing methods still face two major challenges: achieving comprehensive coverage of fine-grained video content, and maintaining factual grounding and global semantic coherence as captions grow longer and more detailed. To this end, we propose Observation-Adaptive Video Captioner (OACap), an agentic model that learns to think with video. Rather than relying on a fixed perceptual context, OACap treats the evolving caption itself as the state of reasoning: it identifies missing, unsupported, or semantically inconsistent content, and adaptively acquires additional visual, temporal, or external evidence to verify and refine the description.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.