acceptodds
Under review as a conference paper at ICLR 2027

LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment

Abstract

Zero-shot visual recognition often depends on localized parts, attributes, and textures that are diluted in global image representations. Existing training-free localized methods alleviate this limitation by evaluating multiple image regions, but their candidate locations are typically generated before class semantics determine where to look, restricting language to post-hoc scoring of already sampled evidence. We introduce LAGO (LAnguage-Guided adaptive Object-region focus), which reframes localized recognition as language-guided directed region discovery. LAGO first establishes a compact, class-agnostic object-centric visual initialization, then constructs a soft semantic prototype from the intermediate prediction to actively steer subsequent search toward class-relevant evidence during inference. This object-first, language-second design addresses the circular dependence between recognizing a class and locating its supporting evidence, while preserving complementary local, contextual, and global cues. Across standard and distribution-shift benchmarks, LAGO transfers with a single fixed configuration and uses candidate regions more effectively under matched budgets. Controlled experiments further isolate the proposed mechanism, showing that allowing language to alter the search trajectory retrieves more relevant regions and improves recognition over applying the same semantic signal only after candidate generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.