acceptodds
Under review as a conference paper at ICLR 2027

VideoPLUM: A Large Multimodal Model for Part-Level Visual Grounding in Videos

Abstract

Objects are composed of fine-grained parts, and identifying these parts can enable more precise visual reasoning and grounded language generation. Extending such fine-grained grounding to videos, however, requires identifying and consistently tracking object parts over time. Existing image-based segmentation and grounding methods do not readily address this temporal dimension at the part-level. We introduce VideoPLUM, a large multimodal model for fine-grained, part-level visual grounding in videos. Our approach combines the semantic reasoning capabilities of a vision-language model with SAM-based segmentation to generate accurate spatiotemporal masks for object parts, proposing an end-to-end approach for fine-grained, part-level grounding across video frames. To address the scarcity of fine-grained part-level grounding annotations in videos, we develop an automatic data-generation pipeline that identifies relevant object parts and temporally propagates their masks. We evaluate our model on both part-level and object-level grounding across videos and images. VideoPLUM substantially outperforms existing baselines in fine-grained part grounding, achieving 51.3 J&F across SMITE, PUMaVOS, and EgoFun3D and surpassing GLaMM + SAM2 by 12.5 points. Our method also generalizes comparably to image segmentation, extending its fine-grained visual reasoning capabilities beyond video grounding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.