FastVLM2: Efficient On-Device Dense Grounding for Generalist VLMs
Abstract
Generalist vision-language models (VLMs) with dense grounding capabilities can perform a broad range of classical vision tasks through a single model and natural-language interface. However, efficient on-device inference with such models faces two complementary bottlenecks: costly high-resolution visual encoding and lengthy autoregressive decoding of numerous spatial locations. We introduce FastVLM2, a unified framework that addresses both. For input-side efficiency, we introduce a native-resolution strategy that adapts pretrained convolutional and hybrid convolution-transformer backbones (ConvNeXt and FastViTHD, respectively) through VLM instruction tuning, enabling arbitrary image resolutions and aspect ratios without additional vision pretraining, image tiling, or explicit layout tokens. On a MacBook with an M1 Max, FastVLM2-4B achieves VQA accuracy comparable to Qwen3-VL-4B with lower prefill latency. For output-side efficiency, we use quantized location tokens to compactly represent grounding outputs, reducing the output-token budget by over relative to textual coordinate representations. To recover the localization accuracy lost to quantization, we introduce a simple Gaussian distance-aware soft cross-entropy objective that brings quantized grounding close to textual-coordinate performance while preserving this decoding advantage. At higher grounding accuracy, FastVLM2-1.8B achieves lower end-to-end latency than Rex-Omni. We further introduce a training-free, geometry-aware box confidence estimator that directly exploits the predictive distributions of quantized location tokens. We show that this scoring rule generalizes across three independently trained grounding models and consistently improves COCO mAP by more than points over the score-free ranking protocol used in Rex-Omni. Together, FastVLM2 makes accurate, confidence-aware dense grounding practical for on-device generalist VLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.