Projective Geometry for Generalizable Multi-View Pedestrian Detection
Abstract
Multi-view pedestrian detection (MVPD) seeks to localize pedestrians on the ground plane by gathering observations from multiple cameras. Recent MVPD methods achieve strong in-domain performance by projecting per-view features onto a shared Bird's-Eye-View (BEV) plane, where localization is performed. Since BEV construction couples learned representation with camera geometry, their spatial reasoning becomes strongly biased toward scene-specific occupancy patterns, limiting generalization to unseen environments. Rather than expanding the training distribution, we address generalization by integrating projective geometry into a scene-agnostic framework. Specifically, each pedestrian is represented as a ground-anchored 3D quadric and its 2D representation in image view is modeled as a projected 2D conic, connected by an exact, invertible () mapping under the pinhole model. As 3D localization is fully lifted from multi-view 2D predictions through camera parameters, the framework learns no scene-specific priors and naturally inherits the cross-scene transferability of pretrained image detectors. Based on this hypothesis, we propose **PGMVG** , a query-based framework that iteratively refines per-view detections and yields pedestrian locations through constrained projective triangulation. Extensive experiments across multiple benchmarks show that PGMVG improves MODA by on cross-scene and on sim-to-real settings, while maintaining strong in-domain performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.