SD-LLaVA: Dual-Stream Vision-Language Reasoning with Structured Semantic Priors for BEV Understanding
Abstract
Large vision-language models (LVLMs) remain unreliable on Bird’s-Eye-View (BEV) semantic maps because their visual representations are primarily optimized for natural-image appearance, while BEV semantics are encoded through discrete labels, spatial structures, and geometric conventions. To address this representation mismatch, we propose SD-LLaVA, a dual-stream architecture that combines continuous visual features with explicit structured semantic priors for BEV-grounded reasoning. A SemanticPriorEncoder extracts class-aware spatial statistics from BEV maps and projects them into semantic tokens aligned with the language-model embedding space. These tokens are jointly integrated with CLIP patch representations, enabling the model to capture complementary visual context and explicit semantic structure. We further introduce a BEV question-answering protocol that separates factual grounding from open-ended driving-oriented reasoning, allowing systematic evaluation of both semantic understanding and response quality. Experiments on a BEVCar-derived benchmark show that SD-LLaVA achieves 90.97% semantic accuracy, substantially outperforming CLIP-only and semantic-only variants, while also improving subjective driving-oriented QA. These results demonstrate that bridging continuous visual representations with structured semantic priors provides an effective approach to grounding LVLMs in symbolic spatial environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.