acceptodds
Under review as a conference paper at ICLR 2027

Recovering and Exploiting Latent Presentation Preferences in Large Language Models

Abstract

LLM-based search, question-answering, and recommendation systems must select among multiple, potentially conflicting external sources. This creates an attack surface: a small amount of attacker-controlled content can disproportionately steer model outputs through changes in presentation, even when the underlying claim is fixed. Existing content-optimization attacks rely mainly on end-to-end outcomes that conflate attention allocation, internal evaluation, and source competition, providing coarse guidance for optimization. We introduce a preference-guided adversarial content optimization framework that recovers model-specific presentation preferences from claim-level Information Attention and sparse Acceptance Vectors and uses these preferences to guide downstream attacks while preserving the core claim and linguistic naturalness. Profiling 120 presentation forms in 11 categories across seven models reveals structured, model-specific preferences. Preference scores correlate strongly with downstream attack success (–). Causal interventions further demonstrate that attention and acceptance components influence evidence adoption. In retrieved-evidence poisoning, our attacks outperform the strongest baselines by 8.67–31.87 percentage points with one controlled source competing against nine sources supporting the correct answer. They also achieve the best target ranks in generative recommendation manipulation, with mean-rank gains of 0.27–2.32 positions. The attacks transfer to black-box models, exceeding the strongest baselines by 23.06–41.03 percentage points in poisoning success rate. These results establish latent presentation preference as a direct and effective signal for optimizing adversarial content to increase the model's acceptance of attacker-controlled claims.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.