GroundingV2X: Grounded Pre-Training for Model-, Task-, and Alignment-Agnostic Collaborative Perception
Abstract
Collaborative perception aims to enhance the perception capability of a single vehicle by enabling information sharing through Vehicle-to-Everything (V2X) communication. However, the heterogeneity of agents poses a significant challenge to the scalability of collaborative perception systems. Existing approaches rely on encoder retraining, feature interpretation, or protocol adaptation to align heterogeneous representations, which are difficult to implement in real-world practice. To address this, we introduce GroundingV2X, a model-, task-, and alignment-agnostic collaborative perception framework via grounded pre-training. Specifically, agents communicate through grounded text derived from their perception results and composed of positional, geometric, and semantic key-value pairs with numerical attributes, providing a unified interface for heterogeneous information exchange. Moreover, we incorporate Fourier numerical embeddings and sub-sentence level attention to preserve numerical precision and entity-specific context with a lightweight text encoder. Then, a position-aware text-to-feature enhancer selectively integrates the grounded text into spatially relevant local features for collaborative perception. In addition, each agent is pre-trained independently using its perceived grounded text, synthetic grounded text, and annotation-derived pseudo-labels for supervision, improving its generalization to diverse textual descriptions. Extensive experiments on OPV2V, DAIR-V2X, and V2V4Real demonstrate that GroundingV2X enables efficient heterogeneous collaboration solely through pre-training while achieving superior performance compared to state-of-the-art methods both in single-task and multi-task scenarios with substantially lower communication bandwidth requirements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.