VideoInspector: Agentic Multi-Video Reasoning with Compositional Segment Training
Abstract
Multimodal large language models have made substantial progress in video understanding, yet most are trained on single-video instruction data and struggle with multi-video reasoning. Such reasoning requires models to retrieve relevant videos, localize crucial moments, and aggregate evidence across different sources. However, these intermediate capabilities are rarely explicitly supervised by existing video training data, leading models to rely on coarse representations rather than grounded cross-video evidence. We introduce VideoInspector, an agentic framework for multi-video reasoning with tool-augmented segment inspection. Given a set of videos, VideoInspector performs coarse-to-fine reasoning by selecting relevant candidates, inspecting fine-grained temporal segments, and composing grounded evidence for final decisions. To train this model, we construct VideoInspector-Instruct, a large-scale segment-level compositional multi-video dataset. Starting from self-collected videos with dense temporal descriptions, we decompose videos into reusable semantic segments and recombine them into multi-video instances across four tasks: cross-video retrieval, cross-video temporal grounding, clip ordering, and video comparison. We further annotate tool-augmented reasoning trajectories and adopt a two-stage training pipeline: cold-start supervised fine-tuning teaches the model to imitate structured reasoning and tool use, while agentic reinforcement learning refines its policy for retrieval, inspection, and termination through exploration. Experiments on multi-video reasoning benchmarks show that VideoInspector improves over its non-tool variant on CrossVid, particularly in candidate analysis and temporal understanding. These results suggest that compositional segment supervision and tool-augmented reasoning provide a promising direction for building more grounded multi-video understanding systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.