CRTrack-Cap: A New Benchmark for Cross-View Dense Video Object Captioning
Abstract
Dense video object captioning (DVOC) localizes, tracks, and describes object trajectories within a video. Existing methods process each camera view independently, producing separate camera-local tracklets and captions for the same person across cameras. These outputs lack a unified description and leave complementary appearance evidence across views unused. We therefore formulate cross-view DVOC (CVDVOC), which associates these tracklets into scene-level global trajectories and generates one caption for each trajectory from synchronized views. To support this task, we construct CRTrack-Cap from CRTrack by auditing text-identity links and converting usable referring expressions into canonical captions. It retains 425 global trajectories and 838 local tracklets across 13 scenes, with canonical captions for 335 identities. We adapt CHOTA to the cross-view setting as CV-CHOTA to jointly evaluate tracking and caption quality. Our reference framework combines View-Exclusive Cross-Camera Association (VECA) for grouping tracklets with Conflict-Evidence Gating (CEG) for integrating caption evidence using its reliability and disagreement across views. Against three single-view methods adapted with MvMHAT for cross-view association, our framework achieves CV-CHOTA scores of 83.43 and 33.48 on the scene-disjoint in-domain and cross-domain tests, respectively. Together, CRTrack-Cap and our framework provide a benchmark and a strong baseline for describing the same person consistently across camera views.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.