Lap3R: Large-Scale Panoptic 3D Reconstruction from Aerial Imagery
Abstract
Large-scale panoptic 3D reconstruction from aerial imagery requires globally consistent geometry and persistent object identities across distant views. Existing feed-forward methods struggle with both: depth inconsistencies distort local reconstructions, repetitive layouts confound loop closure, and cross-view instance supervision is scarce. We present Lap3R, a framework that combines semantic-prior geometry rectification, spatially and temporally verified loop mining, and cross-view co-visibility mask linking. Lap3R first rectifies local geometric distortions using semantic priors. It then verifies long-range loop candidates to retain genuine revisits while rejecting false matches caused by repetitive aerial layouts. To establish persistent instance identities, Lap3R aggregates mask agreement across co-visible views, without temporal tracking or manual cross-view annotations. On large scenes, Lap3R achieves the lowest trajectory error and Chamfer distance among the evaluated methods, while substantially improving semantic and instance reconstruction over feed-forward baselines. We further use Lap3R to generate pseudo-labels from aerial videos; a model trained solely on these labels surpasses the evaluated panoptic baselines on all reported metrics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.