VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
Abstract
As Vision-Language Models (VLMs) rapidly advance toward physical deployment, the predominant focus has remained on action-oriented Embodied AI evaluated on subject-centric consumer video. This trajectory overlooks a pervasive class of Physical AI: Infrastructure AI, which relies on fixed-infrastructure cameras for open-loop insights like safety monitoring and operational logging. How well VLMs process dense, fixed-viewpoint visual data remains insufficiently explored. We introduce VANTAGE-Bench, an evaluation benchmark that measures this "Infrastructure AI Gap," distinguished by four features: 1) Domain Relevance, spanning three operational environments (Warehouse, Transportation, and Smart Spaces); 2) Modality Breadth, unifying image and video evaluation to probe semantic, spatial, temporal, and spatio-temporal capabilities; 3) Task Format Diversity, moving beyond multiple-choice to eight task formulations spanning discriminative QA, generative dense captioning, and spatio-temporal grounding; and 4) Evaluation Novelty, a single-pass trajectory-generation protocol for Single Object Tracking and, to our knowledge, the first such evaluation on fixed-camera infrastructure video, scored against specialist trackers. Annotation spans three regimes over 3,346 expert-annotated media assets: video tasks (3,342 annotations over 854 videos), image grounding (4,281 over 1,864 images), and dense detection (27,404 boxes over 628 images). Evaluating 18 models zero-shot, we find the shortfall relative to consumer-centric benchmarks is concentrated rather than general. Event verification and temporal localization fall 11.7 to 23.5 points at all three scales we can match, while video question answering stays within 5.3 points of VideoMME and 2D spatial pointing is flat or better against BLINK. The temporal pillar is the weakest: GPT-6 Astra is the only model of the eighteen to clear 55.71 mIoU on temporal localization or 37.28 SODA on dense video captioning, leading the field by 9.4 mIoU on the former and 0.9 SODA on the latter. On tracking, the strongest frontier model matches purpose-built specialist trackers at every horizon we test, while weaker open-weight models fall to a static-box baseline. Open-weight models lead 2D object localization outright, so neither parameter count nor proprietary access accounts for the pattern. The dataset, evaluation harness, and public leaderboard are publicly released; the harness is available at an anonymized repository, and the dataset and leaderboard URLs are withheld for double-blind review.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.