Locate in Any Scene: Agentic Object Localization with Multimodal Large Language Models under Diverse Visual Degradations
Abstract
Object localization in real-world scenarios remains challenging due to diverse image degradations and heterogeneous object distributions, which substantially limit the generalization of existing grounding models. Conventional approaches, including scene-specific representation learning and end-to-end pipeline design, are inherently constrained by predefined conditions and therefore lack the flexibility required for evolving environments. In this paper, we propose LocAS, an agentic framework that formulates object localization as an adaptive decision-making process. Rather than relying on static pipelines, LocAS employs a Multimodal Large Language Model (MLLM) as a central agent to dynamically compose localization workflows by selecting from a toolbox of restoration modules and specialized detectors. Specifically, LocAS consists of two key components: Self-Adaptive Image Restoration, which determines whether and how an input image should be enhanced for downstream detection, and Multi-Expertise Localization, which coordinates multiple domain-specialized detectors and consolidates their predictions through instance-level reasoning. To further improve decision quality under fine-grained conditions, we introduce Experience-Aware Calibration and extend LocAS to LocAS-X. LocAS-X accumulates node-level decision experience from a small set of annotated samples and incorporates it into agent memory to support more informed reasoning during inference. This mechanism enables the system to progressively refine its decision policy and better adapt to diverse and dynamically changing degradation scenarios. Extensive experiments on eight challenging benchmarks demonstrate that LocAS-X significantly outperforms existing MLLM-based grounding models, achieving an average improvement of 30.59% in F1 score. These results highlight the potential of agentic localization and provide a promising foundation for robust object localization in complex and dynamic real-world environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.