VISA: VLM-Guided Instance Semantic Auditing for 3D Semantic Occupancy Prediction
Abstract
Semantic 3D occupancy provides a compact scene representation for autonomous driving and robotic decision making, yet closed-set object semantics remain challenging for rare and visually confusable classes. A common way to transfer vision-language model (VLM) knowledge is to align 3D features with language embeddings; however, we find that such alignment can improve text-space objectives without reliably improving voxel-level semantic prediction. We introduce VISA, a task-aligned VLM supervision framework that uses an offline VLM as a structured semantic auditor rather than an embedding target. For each physical object instance, VISA produces a reliability-aware audit containing a closed-set class hypothesis, plausible confusions, and visual attributes, propagates the audit along the object track, and grounds it only to matched 3D object voxels. The resulting structured supervision is distilled directly into semantic logits, while the underlying occupancy architecture and inference pipeline remain unchanged. On nuScenes, VISA yields 3.75–5.72% relative mIoU improvements across two backbones, with 6.16–9.92% relative object-mIoU gains and 8.46–11.70% relative rare-class gains. Beyond autonomous driving, VISA achieves 1.47–2.68% relative mIoU improvements on NYUv2 across four monocular SSC architectures, extending its benefits to indoor scene completion from single RGB images without video input. These results support the applicability of VISA across outdoor and indoor settings, covering both temporal occupancy models and single-image scene-completion architectures.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.