PostCrop: Compose Before Crop for Context-Aware Fine-Grained Visual Perception
Abstract
Large vision-language models (LVLMs) have demonstrated strong fine-grained perception of relations among multiple entities by controlling visual context—expanding or reducing the surrounding context of multiple entities (e.g., zoom-out or locate-then-zoom-in). While existing training-free methods typically rely on either a coarse-to-fine search paradigm or a Crop Before Compose paradigm that crops entity regions before composing them with context, both face a context paradox: crops containing only the entity regions can obscure relations among entities, whereas crops keeping context introduce visual interference. To characterize this paradox, we fix the entity regions and vary the kept context. Across HR-Bench-4K and HR-Bench-8K, we find that the preferred amount of context varies across settings, motivating adaptive context selection. Further analysis shows that the early separation of entity regions in the Crop Before Compose paradigm makes this adaptation difficult. Leveraging these insights, we propose PostCrop, a training-free framework following the Compose Before Crop paradigm that considers relation-relevant context before the final crop is formed. The Evidence Localization and Composition module localizes the region that best matches each queried entity and composes the cue maps of multiple entities on the original visual grid to identify surrounding regions that support the queried relation. The Evidence Cropping and Answering module uses the composed map to determine how much surrounding context to retain and selects a compact crop containing the entity regions and the kept context, which is combined with the original image for answering. Extensive experiments on VSTAR, HR-Bench-4K, and HR-Bench-8K demonstrate that PostCrop consistently outperforms the compared baselines across all three evaluated LVLMs (e.g., 74.0% accuracy on the Fine-grained Cross-instance Perception (FCP) subtask of HR-Bench-4K with Qwen2.5-VL-7B), validating the effectiveness of the Compose Before Crop paradigm.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.