Beyond Fixed Sampling: Uncertainty-Aware Retrieval and Spatially-Grounded Reasoning for Fine-Grained Visual Recognition
Abstract
Recent fine-grained visual recognition methods increasingly adopt a retrieve-then-reason paradigm powered by multimodal large language models. However, this paradigm is often hindered by two critical bottlenecks: Rigid Selection Inflexibility and Ungrounded Speculation. Specifically, conventional top- retrieval fails to adapt to varying sample difficulties, frequently introducing irrelevant noise or missing crucial candidates. Furthermore, holistic image processing during reasoning lacks fine-grained spatial grounding, which often triggers hallucinations and yields unverified visual evidence. To overcome these issues, we propose \ours, a framework designed to refine both the retrieval and reasoning trajectories. First, we introduce Dynamic-, an uncertainty-aware candidate selection mechanism that adapts the probability budget to sample-level confidence and determines the candidate set through cumulative-probability selection. Second, we develop a Discriminative Attribute Anchor (DAA) mechanism for the reasoning phase. By synergizing spatially-grounded evidence with decoupled class attributes, DAA guides the model to anchor onto discriminative local regions while filtering out non-essential semantic features. Comprehensive experiments demonstrate that \ours establishes new state-of-the-art benchmarks across 11 fine-grained datasets, achieving a significant average accuracy improvement of .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.