acceptodds
Under review as a conference paper at ICLR 2027

How To Detect Encoded Reasoning in LLMs? A Decision-Theoretic View of Steganography

Abstract

Large language models are beginning to show steganographic capabilities, which could allow misaligned models to evade oversight mechanisms such as chain-of-thought monitoring. Yet principled methods to detect and quantify such behaviours are lacking. Classical definitions of steganography, and detection methods based on them, require a known reference distribution of non-steganographic signals. For the case of steganographic reasoning in LLMs, such a reference distribution generally unavailable, rendering these approaches inapplicable. We propose an alternative, **decision-theoretic view of steganography**. Our central insight is that steganography creates an asymmetry in usable information between agents who can and cannot decode the hidden content, and this otherwise latent asymmetry can be inferred from the agents' observable actions. To formalise this perspective, we introduce generalised -information: a utilitarian framework for measuring the amount of usable information within some input. Building on this, we define the **steganographic gap** as the difference between the usable information a signal provides to the audited agent (the Receiver) and to a trusted reference agent (the Sentinel), allowing to quantify the amount of encoded information without requiring a reference distribution. We empirically validate our formalism, and show that across different controlled setting it successfully detects and quantifies various encoded reasoning schemes, including semantic ones missed by LLM judges and paraphrasing-based detectors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.