acceptodds
Under review as a conference paper at ICLR 2027

RSHR-Bench: Benchmarking MLLM Understanding and Visual-Token Pruning in Ultra-High-Resolution Remote Sensing

Abstract

Multimodal large language models (MLLMs) have made substantial progress in visual perception and reasoning. However, reliable evaluation of MLLMs on ultra-high-resolution remote sensing imagery remains limited because existing high-resolution benchmarks do not consistently include systematic human verification of visual grounding. To address this gap, we introduce the Remote Sensing High-Resolution Benchmark (RSHR-Bench), which contains 5329 satellite, aerial, and UAV images with an average resolution of \(8700*8065\) pixels. RSHR-Bench comprises 1932 perception and reasoning VQA items, spanning multiple-choice and open-ended formats as well as single-image, multi-image, and multi-turn settings. All questions and reference answers are independently verified by human annotators for correctness, visual grounding, and unambiguous wording. We evaluate thirty-four vision–language models on RSHR-Bench, ranging from general-purpose MLLMs to models developed specifically for remote sensing. The results show that the highest overall accuracy is only 54.1% and that most models specialized for remote sensing fail to outperform general-purpose MLLMs. We further compare twenty-two visual-token pruning methods on RSHR-Bench, XLRS-Bench-lite, and MME-RS at multiple nominal token-retention budgets. Under aggressive compression, many methods suffer substantial accuracy losses, highlighting the trade-off between token reduction and the preservation of visual understanding. We hope RSHR-Bench will serve as a challenging benchmark for advancing MLLMs on ultra-high-resolution remote sensing imagery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.