acceptodds
Under review as a conference paper at ICLR 2027

OpenSight: Explainable Open-Set Fine-Grained Recognition with MLLMs

Abstract

Multimodal large language models (MLLMs) have demonstrated strong general capabilities. Still, they remain unreliable in open-set fine-grained visual recognition, where the correct decision is sometimes to reject all provided candidate categories. These tasks often rely on subtle visual distinctions to tell different categories apart. Through controlled analysis, we first find that perception is not the dominant reason for failed recognition. After extracting visual features, models often make decisions based on memorized category priors rather than grounded visual cues, leading to unfaithful reasoning. To this end, we introduce OpenSight, a reasoning-guided reinforcement-learning pipeline that curates samples with verified, attribute-grounded reasoning rationales to enhance fine-grained reasoning in MLLMs. By explicitly identifying hard negatives via subtle visual contradictions, OpenSight enables the model to ground its textual explanations in specific visual evidence. Importantly, the acquired reasoning pattern can be generalized to domains that are visually distinct from the training domain. Across six fine-grained open-set benchmarks, our approach improves the average accuracy by 4.1% and 13.0%, and the harmonic mean by 16.8% and 13.6% over standard GRPO-tuned and base Qwen2.5-VL, respectively. Qualitatively, we show that OpenSight facilitates more evidence-linked explanations and more reliable abstentions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.