UAVOcc: Hierarchical Spatiotemporal Semantic Occupancy Prediction for UAVs
Abstract
Aerial semantic occupancy faces a resolution–coverage trade-off under a fixed voxel budget: fine discretization preserves local geometry but restricts physical extent, while coarse discretization expands coverage at the expense of detail. We formulate hierarchical aerial semantic occupancy, representing these competing requirements with complementary fine-local L1 and coarse large-range L2 physical levels of distinct metric resolutions and extents. To make this formulation measurable, we establish a benchmark spanning three virtual and real-world aerial data sources, with a common L1/L2 task definition and Package-Physical evaluation for non-redundant joint physical coverage. Modeling this hierarchy requires cross-level interaction beyond ordinary multi-scale fusion, as the two levels differ in metric correspondence and temporal support. We therefore develop UAVOcc with Global–Local Cross-Level Interaction (GLCI), which combines metric-local 3D correspondence with broader global context through bidirectional spatial exchange; pose-aligned same-level temporal fusion is further followed by coarse-to-fine Temporal GLCI to exploit the broader historical support of L2. Across all three benchmarks, UAVOcc consistently outperforms adapted occupancy baselines, improving Package-Physical mIoU by 3.41, 5.47, and 5.27 points over the strongest listed baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.