SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
Abstract
Multimodal search agents must reason over long horizons, yet existing pipelines design training data, search environments, and rewards in isolation. This fragmentation discards synthesis metadata, relies on irreproducible external engines,and leaves RL with sparse trajectory-level supervision. We introduce SearchEyes,a knowledge-graph-based simulated search world that unifies these components.Perception-Knowledge Chains (PKC) sample constrained multi-hop paths over visual and knowledge entities while retaining hop-level metadata that defines both a self-contained environment and reward anchors. Hop-Anchored Policy Optimization (HaPO) uses these anchors for step-level credit assignment without a learned process reward model. Across six multimodal knowledge-intensive benchmarks, SearchEyes achieves state-of-the-art performance among opensource multimodal search agents; SearchEyes-27B improves over the strongest open-source baseline by 6.2 points on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.