acceptodds
Under review as a conference paper at ICLR 2027

Where Does Occupancy Latency Go? A Per-Stage Decomposition of 3D Occupancy Inference

Abstract

Evaluation of 3D semantic occupancy prediction relies almost exclusively on accuracy metrics, namely mean intersection over union (mIoU) and its ray-based variants, measured on server-grade GPUs. Deployment-related metrics do appear in the literature, but they are collected on an ad hoc basis across different hardware platforms under unspecified settings. Among the works we surveyed, none reports these metrics as a function of computational budget. To address this gap, we provide the measurements needed for budget-aware evaluation by decomposing occupancy inference latency stage by stage. We instrument six public configurations and attribute latency to six concrete stages, repeat runs under a common fixed protocol, report median and 95th-percentile (p95) latency, and account for the residual time outside the measured stages. Extensive experiments show that the 2D image backbone, rather than the occupancy head, dominates single-frame latency. A two-frame temporal window shifts the dominant cost to view transformation, while scaling up the backbone makes image feature extraction a co-dominant cost again. Comparisons among variants within the same architecture lineage further show that the differences are concentrated primarily in the bird's-eye-view (BEV) encoder backbone and occupancy head, where they correspond to approximately constant absolute latency savings. We use this decomposition to introduce a new reporting convention that expresses accuracy as an explicit function of deployment budget. We apply it to Occ3D-nuScenes and make the measurement framework and results publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.