Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification
Abstract
Spatio-temporal video grounding (STVG) requires a model to identify when a queried event occurs and track the referred target throughout that interval. Applying multimodal large language models (MLLMs) frame by frame produces long visual contexts and autoregressive coordinate sequences. We formulate MLLM-based STVG around second-level trajectory control points and reconstruct dense tracks with deterministic inter-second smoothing. Training combines continued pre-training on broad grounding data, supervised fine-tuning with generated decision traces whose numerical labels are replaced by ground truth, and reinforcement learning with a verifier based on temporal and motion-aware spatial overlap. On 10-second videos, the second-level protocol reduces end-to-end latency by and prefill compute by relative to a 25-FPS frame-level alternative; on 60-second videos it remains executable when 5-FPS and 25-FPS inputs exceed the context window. A 9B task-aligned model improves over its SFT checkpoint on VidSTG and HC-STVG and also improves its base checkpoint on Video-MME and MMVU, demonstrating that compact grounding can coexist with general video understanding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.