acceptodds
Under review as a conference paper at ICLR 2027

HerBench: Understanding Fine-Grained Visual Recognition in Multimodal Agents through Trait-Level Evaluation

Abstract

Multimodal Large Language Models (MLLMs) equipped with active visual tools have significantly advanced fine-grained visual recognition (FGVR) by inspecting local image regions. However, despite access to high-resolution visual evidence, multimodal agents still struggle with reliable specimen identification, leaving a fundamental question unanswered: *Do identification failures arise from selecting the wrong traits, misjudging subtle trait states, or failing to use those states to distinguish candidate categories?* To dissect this visual reasoning chain, we formulate a trait-level evaluation across three dimensions: *trait selection*, *trait judgment*, and *the use of trait evidence for identification*. We introduce **HerBench**, a specialized benchmark linking 851 megapixel-scale herbarium specimens (spanning 239 species) to 5,325 fine-grained trait questions derived from taxonomist-authored dichotomous keys (dKeys). Furthermore, inspired by professional taxonomic workflows, we propose a dKey-guided agentic framework that interleaves active visual inspection with expert key traversal to systematically guide trait selection and evidence integration. Extensive evaluations across nine state-of-the-art multimodal agents reveal a substantial *trait selection gap*: agents frequently answer isolated trait questions correctly yet completely omit or fail to make definite judgments on the same traits during end-to-end identification. Moreover, even when observation targets are explicitly specified, trait judgment accuracy remains modest (25.51%–64.55%). While dKey guidance improves species F1 for eight of nine models, controlled experiments with reference traits demonstrate that its performance gains are substantially amplified when accurate trait states are supplied, highlighting that the true utility of expert guidance depends fundamentally on the fidelity of the agent's perceptual evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.