acceptodds
Under review as a conference paper at ICLR 2027

RL-Trained Crop-and-Zoom Is Often Decorative: Pixel Budgets, Protocol Costs, and Causal Credit

Abstract

We ask when crop-and-zoom tools trained with reinforcement learning (RL) help and why they often fail. Many tool-equipped vision-language models (VLMs) turn out to behave decoratively: they localize accurately, yet the returned crop adds little causal information. The cause is a pixel-budget gate: a perfect crop's value is set by the gap between an image's resolution and the model's working pixel budget, and in every configuration we measure grows strictly monotonically as that budget tightens (from +5.8 to +31.1 points across datasets and scales). Real above-budget imagery (MME-RealWorld-Lite, DocVQA) confirms the degradation the gate rests on: whole-image accuracy falls under a tight budget only for models that can use the pixels. Standard benchmarks and RL training pools (median 0.19M pixels) sit below that budget, so crops that are useless there carry recoverable information at deployment budgets. Where a crop does carry value, the standard additive tool-call serialization throws it away: a plain second image helps in all 16 Qwen dataset-by-budget comparisons (+6.8 to +17.3 points), but routing the identical pair through a tool-call exchange costs the untrained 3B base model 16–52 points across our budget sweeps, and larger untrained models 4–25 at 12.8M, by inducing tool-call loops; RL training only reduces that loss without learning to use the crop. Letting the returned crop replace the full image, or presenting it plainly, recovers most of the loss with no retraining. We audit tool use with counterfactual ablation credit: a call counts only if replacing its returned crop with a shifted crop flips the answer. On V* it is empirically nested within probe-style credit (p < 1e-4); yet, under the standard group relative policy optimization (GRPO) recipe, even content-causal credit with above-budget data does not yield durable, selective tool use. Together, these results reduce to three conditions governing when crop-and-zoom can pay: images served above the model's working budget, the crop presented in place of the full image, and content-causal credit.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.