acceptodds
Under review as a conference paper at ICLR 2027

Promotion Under Audit: Subliminal Manipulation of Interactive LLM Advertising

Abstract

With interactive large language model (LLM) assistants increasingly mediating consumer decisions, commercial incentives arise for model providers to steer recommendations toward sponsored options. However, as LLM-generated content faces increasing scrutiny, advertising success must be assessed not only by its capability to steer recommendations toward sponsored options, but also by whether such steering remains effective under audit. Existing manipulation methods always leave detectable evidence through biased instructions, anomalous triggers, or formulaic promotional wording. We propose Subliminal-On Policy Distillation (OPD) to eliminate observable evidence, which transfers a commercial preference into model weights using OPD on semantically unrelated training data. The resulting model requires no biased prompt, Retrieval-Augmented Generation (RAG) injection, or trigger at deployment. Jointly accounting for promotion success and full-context detectability, Subliminal-OPD achieves the best performance on controlled clinic recommendation tasks. Complementary judge-independent analysis of successful responses in sentence-embedding space shows that Subliminal-OPD induces the nearly smallest distributional shift from the unmanipulated model, providing further evidence of audit resistance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.