LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling
Abstract
Rapid advances in the agentic capabilities of large language models have spurred the development of increasingly challenging benchmarks for deep search agents. Constructing high-quality questions for these benchmarks often relies on labor-intensive expert authoring to ensure factual validity and sufficient difficulty. To automate this process, we introduce LoHoSearch (Long-Horizon Search Agents), a benchmark comprising 544 automatically verified and human-reviewed questions across seven domains. Its automated pipeline samples tree- and graph-structured subgraphs from a Wikipedia-derived knowledge graph (KG) containing 7.62 million entities. The KG enables systematic control over candidate-set sizes and constraint structures to shape question difficulty, while supporting answer uniqueness checks across the full graph to help ensure question validity. The sampled subgraphs are then converted into natural-language questions, which undergo automated validation, difficulty filtering, and human review. Under a standardized evaluation setup, even the best-performing model achieves only 52.02% accuracy. With DeepSeek-V4-Flash, successful trajectories use 1.7× as many tool calls on average as on BrowseComp, highlighting the long-horizon search and reasoning demands of our questions. Existing context-management strategies yield a maximum accuracy gain of 6.8 percentage points, substantially smaller than the 14.03-point gain on BrowseComp. These findings highlight the need for more effective context-management methods. LoHoSearch provides a demanding testbed for evaluating evidence discovery, multi-constraint reasoning, and context management over extended search trajectories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.