acceptodds
Under review as a conference paper at ICLR 2027

RIVET: Evaluating Retrieval-Induced Value Steering and Retrieval-Stage Mitigation in Search-Augmented LLMs

Abstract

Agentic search pipelines decide, unseen by users, which query to issue and which retrieved snippets to deliver. Whoever controls these choices could shift a model's expressed values without altering any document, a threat that corpus-poisoning defenses and standard RAG metrics are not designed to catch. We propose RIVET, a controlled protocol that evaluates this vulnerability by varying only query formulation or evidence selection over real search snippets. Across five LLMs and two value benchmarks, RIVET shows that the risk is real: on items whose retrieved pool holds both sides, selecting which authentic snippets reach the model moves its judgments by up to 23% of the response span, a single snippet suffices, and adding a single word (support or oppose) to the search query also shifts them, most clearly on ValueActionLens. This is not blanket context-following, since settled factual claims move least. Because the steering enters through a context that misrepresents what the model's own query found, we propose M-RIVET, a retrieval-stage defense that repairs a context only when it is one-sided relative to what its own neutral query retrieved. Across three models and both benchmarks, it removes 53–90% of the steering from oppose-oriented selection, remains effective with independent labels, and leaves unflagged contexts unchanged. Retrieval-induced value steering is thus both a real risk and one that can be repaired at the retrieval stage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.