Seeing but Struggling to Say: Organizing Dispersed Evidence in Multi-modal Large Language Models for Fine-Grained Recognition
Abstract
Multi-modal large language models (MLLMs) have shown remarkable general capabilities, yet still struggle with fine-grained visual recognition (FGVR), a long-standing task that demands distinguishing subordinate-level categories through subtle visual cues. While existing efforts focus on improving MLLMs for FGVR by introducing external supervision, we revisit this challenge by asking whether MLLMs already possess fine-grained understanding intrinsically. Through layer-wise linear probing, we find that the fine-grained evidence is indeed present across the hidden states, but can hardly be decoded by the native generation. Furthermore, the evidence disperses away from the answer position at intermediate layers, and this dispersion hurts fine-grained recognition far more than coarse-grained one. What MLLMs lack is thus not the fine-grained knowledge, but its organization into a discriminative structure that the readout can exploit. Motivated by this, we propose **O**rganizing **A**ggregated fine-grained evidence with **S**elf-el**I**cited attribute**S** (OASIS), a simple yet effective method that aggregates the dispersed evidence at the middle layer and organizes it by attributes self-elicited from the MLLM itself. With the backbone entirely frozen, OASIS unlocks the intrinsic fine-grained capability of MLLMs across six benchmarks and four backbones, while mitigating the intermediate-layer dispersion and instilling fine-grained discriminative structure. The code is available at https://anonymous.4open.science/r/OASIS_MLLM.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.