IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models
Abstract
Instance-level recognition (ILR) concerns distinguishing individual instances from one another, with person re-identification as a prominent example. Despite the impressive visual perception capabilities of modern VLMs, we find their performance on ILR unsatisfactory, often dramatically underperforming domain-specific ILR models. This limitation hinders many practical applications of VLMs, e.g., where recognizing familiar people and objects is crucial for effective visual understanding. We attribute this limitation in part to the VLM's general-purpose visual representation, which is not explicitly trained to distinguish visually similar instances. To overcome this challenge, we propose IIR-VLM, a VLM enhanced for In-context Instance-level Recognition. We integrate pre-trained ILR expert models as auxiliary visual encoders to provide highly instance-discriminative features complementary to those of the VLM's general-purpose encoder. Through a lightweight two-stage training process, IIR-VLM aligns the expert encoder with the base model, learns to identify new instances in-context, and leverages this knowledge for instance-aware visual understanding. Experiments confirm the benefit of instance-discriminative visual pretraining, particularly for visually similar instances. On IIR-Bench, our challenging new benchmark spanning varying difficulty levels and diverse categories including persons, faces, pets, and general objects, IIR-VLM substantially outperforms strong VLM baselines and prior specialized methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.