HAE-GEO: Benchmarking Deep-Search Agents under Hierarchical Web Evidence Poisoning
Abstract
Search-augmented LLM agents increasingly mediate consumer decisions, making them targets of Generative Engine Optimization (GEO) poisoning: fabricated reviews, rankings, and endorsements published online so that generative search systems retrieve and trust them. Existing benchmarks measure whether manipulated content is retrieved or endorsed, but not how agents handle poisoned evidence after exposure. We introduce HAE-GEO, a benchmark that holds a fabricated brand and its false claim fixed while escalating how the claim is presented, from direct assertion (L1) to contextual camouflage (L2) and apparent corroboration across multiple pages (L3), within a controlled corpus of 72,039 clean pages and 770 poisoned pages per level. Agents interact through a Search–Scrape interface that separates seeing a search result from reading the full page; each trajectory is mapped to five observable states (exposure, adoption, verification, recovery, endorsement) and scored with a six-dimensional rubric. Across ten agents, we find that (1) poison recognition declines from L1 to L3, and the decline persists even when page type and length are matched—a pattern we call the corroboration trap; (2) agents rarely obtain independent evidence after adopting a fake brand, while defense prompting increases both verification attempts and strict evidence acquisition; (3) iterative search lowers fake-brand endorsement without consistently improving recognition, whereas explicit reasoning improves both; and (4) defense prompting raises false rejection of real but little-known brands from 11.4% to 18.0%. Evaluating Web agents therefore requires examining how they verify evidence, not only what they recommend. HAE-GEO provides a diagnostic testbed for building agents that verify, rather than merely avoid, suspicious evidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.