acceptodds
Under review as a conference paper at ICLR 2027

AgroGround: Multi-Granularity Grounded Recognition in Agriculture

Abstract

Agricultural visual models are generally evaluated for either recognition or localization, but reliable diagnosis requires identifying what is present and localizing the visual evidence that supports the prediction. Agricultural visual question answering (VQA) datasets contain rich semantic labels but often lack annotations linking them to image regions. Adding these annotations by hand is costly at scale. We introduce AgroGround, a large-scale dataset for grounded agricultural recognition, where models identify plant diseases and other agricultural targets and localize their corresponding image regions. We develop an automated annotation pipeline that uses semantic labels from eight existing agricultural VQA datasets to annotate both fine-grained regions such as disease lesions and whole objects, producing 794,850 instruction examples. Healthy images provide negative supervision for disease-related queries, teaching the model to return empty predictions. We train a shared vision-language model through supervised fine-tuning, combining known-target grounding instructions with instructions that require both target recognition and localization. We evaluate predicted identities, regions, their joint correctness, and healthy-image abstention on 1,480 human-verified images disjoint from all training data by file path and perceptual hash. Grounding-only fine-tuning reduces recognition accuracy from 51.8% to 29.1%, and adding recognition-and-localization instructions raises it to 72.6%. With positive images and spatial annotations held fixed, combining the two instruction formats raises joint accuracy from 19.2% to 43.3% with comparable grounding performance. Healthy negatives raise abstention on healthy images to 95.0%, and reinforcement learning improves lesion-level grounding. The resulting 2B model achieves higher grounding F1 than its annotation teacher on both our benchmark and the external PlantSeg test set. AgroGround establishes a benchmark for grounded agricultural recognition, measuring joint correctness of target identity and localization, as well as abstention on healthy images.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.