Choosing a Side is not Enough: Epistemic Transparency Under Knowledge Conflict
Abstract
LLMs are increasingly used in professional knowledge work, where their outputs can influence decisions with real-world consequences. When a model encounters information that contradicts its prior knowledge, it should communicate this conflict transparently to the user. Existing work on knowledge conflicts focuses on which source, parametric or contextual, a model chooses, rather than whether models disclose the presence of the conflict. Moreover, existing work primarily studies multiple-choice questions under simplified counterfactual evidence generation schemes. In this work, we address the first gap by scoring model responses against a taxonomy of epistemic transparency, rather than measuring accuracy alone. We address the second gap by proposing a general framework for generating knowledge conflicts grounded in advanced domain knowledge, which we instantiate in this work as a set of real-world tasks based on counterfactual evidence. These contributions allow us to measure epistemic transparency under realistic knowledge conflicts, as well as the impact of the format and plausibility of the evidence, and the real-world stakes of the task. We find substantial differences in transparency across five evaluated frontier models, with none fully matching our definition of optimal, risk-sensitive epistemic alignment. We hope that our work will ground future discussion of epistemic transparency in real-world applications of LLMs, and help establish a risk framework for epistemic transparency across domains and deployment scenarios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.