Defending against GEO Poisoning with Hierarchical and Explainable Byzantine-Robust Aggregation
Abstract
Large language models (LLMs) increasingly rely on retrieval systems and external tools to obtain up-to-date information beyond their parametric knowledge. This reliance introduces a new attack surface. Knowledge corruption attacks can inject misleading evidence into retrieval pipelines, while Generative Engine Optimization (GEO) poisoning can make adversarial content more likely to be retrieved and used during reasoning. We propose a hierarchical and explainable Byzantine-robust aggregation framework with dynamic multi-source evidence retrieval. Our method uses reinforcement learning-guided retrieval, evaluates the stance and reliability of individual evidence items, and organizes claims in a dynamically expanded argument tree. It then aggregates node scores bottom-up to limit the influence of unreliable or conflicting evidence. Experiments on four datasets and two backbone LLMs, together with component and retrieval-tool ablations, demonstrate the effectiveness of ToE on standard claim-verification benchmarks and under controlled GEO poisoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.