CGDP: Complementarity-Guided Dual-Pretext Learning for Collaborative Perception
Abstract
Collaborative perception enables multiple agents to fuse complementary observations for comprehensive scene understanding, yet learning effective representations typically relies on expensive 3D bounding box annotations. Reconstruction-based pretraining offers a promising alternative by recovering masked geometry from unlabeled multi-agent point clouds, but existing methods treat all observations uniformly during masking and condition reconstruction solely on collaborative inputs, neglecting the varying contributions of different agents and providing no explicit mechanism to distinguish what can be recovered from ego-only perception. To bridge this gap, we propose a Complementarity-Guided Dual-Pretext (CGDP) learning framework that introduces two key designs, namely an ego-relative spatial masking strategy that preserves ego-points while selectively masking non-ego regions, and a dual reconstruction objective that combines complementarity-weighted masked collaborative reconstruction with ego-to-collaborative scene prediction. Both tasks are jointly optimized on a shared encoder-decoder architecture using targets constructed directly from unlabeled point clouds, and only the pretrained encoder is transferred to downstream detection without introducing any additional inference overhead. Extensive experiments on V2X-Real and OPV2V demonstrate that CGDP achieves state-of-the-art 3D detection performance, validating that explicitly modeling observation complementarity during pretraining substantially benefits collaborative perception. The code and model will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.