Can You Make Models Believe False Facts in the Prompt?
Abstract
Making LLMs believe false facts is useful for alignment evaluations and for building model organisms. Previous work explored doing so with Synthetic Document Finetuning (SDF). But SDF requires fine-tuning one new model for each set of injected false beliefs. To enable larger scale SDF-powered alignment evaluations, we attempt to produce a single model that believes any fact placed in a tag in the prompt. To do so, we explore doing prompt-conditioned SDF: fine-tuning a model on synthetic documents covering many facts, with each document conditioned on a fact presented in the prompt. We evaluate how deeply the model believes the facts provided in the prompt using evaluations from prior work and new Jacobian-lens-based evaluations. We find that the belief rate on unseen facts scales roughly log-linearly with the number of training facts, making this a cheap, promising avenue for making models believe false facts. We use RL to reduce the rate at which models verbalize that they are following a false fact in the prompt, on top of prompt-conditioned SDF, this generalizes to held-out evals and gets closer to per-fact SDF on most evaluations. But the reward is scored on one of our belief-depth metrics, which makes it harder to trust those evaluations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.