Towards Long-horizon Agentic Multimodal Search
Abstract
Multimodal deep search agents can solve complex tasks by iteratively gathering textual and visual evidence, but long-horizon reasoning is hindered by managing accumulated multimodal contexts while preserving visual evidence for continued use. To address this, we propose LMM-Searcher, a long-horizon multimodal deep search framework built on a file-based visual representation mechanism. By storing visual assets externally and referencing them with lightweight textual identifiers (UIDs), our approach reduces context overhead while preserving multimodal information for future retrieval. We further introduce a data synthesis pipeline for generating complex cross-modal multi-hop reasoning data. Using 12K synthesized trajectories, we fine-tune Qwen3-VL-Thinking-30A3B into a specialized multimodal deep search agent. Experiments on four benchmarks show that LMM-Searcher scales to 100-turn search horizons and achieves state-of-the-art performance among open-source models on challenging benchmarks such as MM-BrowseComp and MMSearch-Plus, while maintaining strong cross-model generalization. Code, data, and models will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.