The Cost of a Near Miss: Compute-Matched Evaluation of Training-Free Visual Cropping
Abstract
Training-free visual cropping aims to improve fine-grained visual question answering, but an appended crop, and the passes that localize it, also add compute. We count visual tokens per question and compare cropping methods to no-crop baselines given the same tokens. ViCrop's reported gain reproduces, and HiDe's keeps its sign at a smaller size. On LLaVA-1.5, which reads images at a fixed resolution, ViCrop's crop beats centered, random or repeated second views of identical token count by 18.9 to 21.5 percentage points on V\*Bench, and two higher-resolution tiles by 13.1. At HiDe's own budget of 16,384 tokens, a baseline allowed all of HiDe's tokens still trails it by 3.4 and 6.5 points on HR-Bench-4K and half of HR-Bench-8K. At low budgets on a backbone that can spend the same tokens on resolution, neither gain survives: ViCrop's once the answering call is matched, HiDe's once its localization is charged. Pooled over 30 Qwen2.5-VL configurations, matching the answering call lowers ViCrop's gain of 1.81 points by 2.65, leaving no gain distinguishable from zero on natural scenes and a loss of 2.03 on documents. Whether cropping pays depends on whether its box finds the target often enough to beat what the baseline gains by spending the same compute on resolution. Annotated-box crops on V\*Bench still gain 19.9 to 30.4 points when matched, but at ViCrop's 256-token setting attention boxes hit the target on only 26.7% to 39.3% of questions. In-place magnification adds no tokens, but pasting an enlarged region over the image can overwrite the target when the box narrowly misses it. Displacing annotated boxes on V\*Bench reveals a near-miss trough: nearby misses score 7.8 to 10.8 points below distant ones at comparable hit rates. Real attention boxes pay this price too. On the 44% of their misses whose pasted patch still covers the target, pasting costs 12.5 points more than on the rest, a gap that resampling the same boxes does not show. Resampling, which compresses the surroundings instead, shows no trough on any of the three models, yet with detector boxes it is significantly worse than pasting on five of six document configurations. Cropping should therefore be compared at matched compute, and in-place methods should report how their primitive responds to localization errors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.