acceptodds
Under review as a conference paper at ICLR 2027

Generalizing VLMs to Extreme Domains: Benchmarking, Reward-Driven Post-Training, and Scaling

Abstract

Modern vision-language models (VLMs) appear general because a single autoregressive interface can solve many visual tasks. Yet task breadth does not imply robustness across visual worlds: current evidence is drawn largely from natural imagery resembling web-scale pretraining data. We investigate whether unified VLM perception survives extreme shifts in image formation and visual statistics. To make this question measurable, we introduce CrossVLM-Bench, a benchmark for both post-training and evaluation across remote sensing, industrial inspection, agriculture, and underwater imagery. It contains 7,362 images and 66,258 image–task pairs, with every image supporting nine tasks spanning recognition, localization, segmentation, grounding, reasoning, and region-level description. Its strictly image-disjoint splits combine human-annotated source geometry with human-verified task and language annotations. Evaluation across four VLM families reveals a systematic breakdown that extends beyond recognition to spatial precision, object-set completeness, and language–region grounding. Target-domain supervised fine-tuning (SFT) partially recovers these capabilities, but optimizes the likelihood of a particular response serialization rather than the correctness of the decoded perceptual outcome. We therefore propose CD-GRPO, a reward-driven post-training framework that recasts extreme-domain adaptation as structured outcome optimization. CD-GRPO verifies complete decoded predictions using task-native semantic and geometric criteria while enforcing valid, complete, and nonredundant output structure. On Qwen3-VL-8B, CD-GRPO raises the average score across four domains and nine tasks from 25.7 to 48.5, exceeding SFT by 6.8 points and the strongest recent perception-oriented RL baseline by 7.7 points under equivalent protocols. Scaling analyses across model capacity, target-domain data, and test-time adaptation further characterize when additional resources translate into domain robustness. Together, our results establish extreme-domain robustness as a distinct axis of VLM generality and structured outcome optimization as an effective route toward it. We will release the benchmark, model, and code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.