A Common Scale for Contextual Override: What Survives a Change of Wording
Abstract
Retrieval systems are vulnerable to attacks embedded within the documents they retrieve. Presenting a claim as fiction, attributing it to a specific source, or repeating it across multiple documents are all known exploit vectors. However, since each finding originates from a different experimental setup, these attacks have never been systematically ranked against one another. Furthermore, they have never been measured against the baseline control required to make such a ranking interpretable: the same claim, embedded in the same document, but explicitly denied. In this work, we aim to construct this unified scale. A controlled document generator maintains the fact, counterfactual object, prompt question, sentence count, and mention count as strict invariants, changing only the framing of a single sentence. The resulting behavior is then scored as a signed margin against each model’s measured parametric prior. Because each rung of this scale uses a distinct sentence, a surface-level ranking of framings risks reducing the analysis to a simple ranking of arbitrary wordings. To mitigate this, we rebuild both evaluation ladders using an independent set of claim templates and re-assess all 11 models. Accumulated corroboration emerges as the most effective attack we tested. All models are overridden on a majority of facts by at least one accumulation condition (achieving an 89% override rate across 185 facts at eight sources, compared to just 26% when the claim is denied). Instruction tuning halves the efficacy of single-mention attacks while simultaneously sharpening the impact of accumulation. This asymmetry is critical for real-world deployment, where a retrieval ranker dictates the exact number of corroborating passages presented to the model. When matched on their baseline parametric priors, models differ primarily in susceptibility (the magnitude by which a single mention shifts their confidence), not in memorization or polarity sensitivity. Ultimately, while the hierarchical ordering between framing tiers survives a change of wording, the assumption that denial constitutes the weakest possible framing does not.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.