Reference-Calibrated Expert Composition for Open-Domain Fine-Grained Recognition
Abstract
A long-standing goal of fine-grained visual recognition is to build a general-purpose system that recognizes fine-grained categories across a broad and growing range of domains. However, the visual cues needed to distinguish closely related categories vary across domains, and models specialized through task- or domain-specific training may perform less reliably outside their target domains. Extending coverage through this approach therefore requires repeated supervision and model adaptation, making broad fine-grained recognition difficult to scale. Visual models pretrained on large-scale data offer an alternative foundation, as their representations already capture rich fine-grained distinctions. Our analysis shows that the examined models have uneven capabilities across domains but exhibit complementary strengths. These observations motivate composing pretrained experts whose strengths differ across domains. We propose a training-free framework that reads domain-specific capability from a small set of labeled reference images and composes heterogeneous frozen experts accordingly. Specifically, references calibrate domain-specific expert contributions for visual comparison and image–name matching, while within-class feature statistics define the visual comparison geometry. Held-out reference predictions further calibrate the balance between visual and semantic evidence. All pretrained models remain frozen, allowing new domains and new experts to be incorporated without retraining. We evaluate the framework on 598 mixed-domain classes, 10,000 species, and 4,431 product identities, demonstrating the utility of reference-calibrated composition across heterogeneous recognition tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.