The System Is Manipulated. Are You? Assessing the Downstream Impact of Manipulated LLM Decision Support
Abstract
LLM alignment is dual-use with implications for advertising, propaganda, and censorship: mechanisms intended to align models with user intent can be repurposed to steer decisions by selectively presenting and reframing evidence. Yet successfully steering an LLM’s recommendations does not necessarily mean successfully steering the downstream decision-maker. This raises a more fundamental question: does LLM decision steering actually work, and can users detect when it is happening? We study such steering in a synthetic RAG decision-support task involving humans and persona-conditioned delegated agents. Our manipulated condition combines poisoned retrieval sources with system-prompt instructions favoring target options. The joint manipulation substantially reduces initial decision accuracy for both humans and delegated agents: by 56.3 percentage points among 64 human participants and by 44.4 percentage points in our main persona-conditioned agent setting. Descriptively, agent runs show more correction among initially incorrect decisions than humans after access to additional evidence (85.2% vs. 28.6%). Detection shows a different pattern: the persona-conditioned agents identify manipulation more often than humans after disclosure (72.8% vs. 46.9%), but also have a substantially higher false-positive rate (59.1% vs. 21.9%). Together, these results show that joint RAG manipulation can bias both human and delegated-agent decisions, with distinct patterns in correction and detection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.