LiveSearchBench: Diagnosing Evidence Access and Use in Multimodal Search
Abstract
An agent can answer a recent visual question without searching, or retrieve the right evidence and still answer incorrectly. End-to-end accuracy alone cannot distinguish these outcomes. We introduce **LiveSearchBench**, a renewable benchmark that connects the construction of search-demanding questions to the diagnosis of search behavior. The collection contains 6,000 questions in 30 dated builds. Evidence-first generation and repeated visual, closed-book, and supplied-evidence checks retain questions whose answers are unexpressed without evidence but recoverable with it under a fixed panel. Evaluation pairs closed-book, with-search, and oracle-evidence runs on the same items, then separates evidence availability from answer correctness. Our experiments show how similar search scores can conceal different evidence profiles, why increased coverage need not improve accuracy, and how admission changes the evaluated population. Together, the construction records, paired conditions, and targeted interventions provide a common basis for asking both whether a search agent succeeds and what information supports that success.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.