acceptodds
Under review as a conference paper at ICLR 2027

Language Models can Learn High-Capacity Secure Steganography

Abstract

Steganography hides messages inside innocuous output so that an observer cannot tell a secret is being sent. This is a safety concern for language models deployed under monitoring. The worst case for a defender is a model whose output is statistically indistinguishable from its normal generation, yet hides as much secret information per token as information theory allows. In this work, we show that language models can be trained to approximate such schemes, and we study the signatures they leave for monitors. We introduce MEC-LLM, an LLM model organism trained using objectives motivated by minimum-entropy coupling (MEC). The trained model encodes uniformly random 128-bit secrets at a raw rate of 0.57 bits per token. A classifier using only text features distinguishes stego from cover with AUC 0.62, while one using the model's full output logits reaches AUC 1.00. We also show that defenders can disrupt the channel by paraphrasing the stegotext with Claude Haiku 4.5: across 13 paraphrase prompts ranging from minimal edits to structural rewrites, bit-error rate rises from 12.8% to 48.7%. Finally, we introduce a complete model organism that pairs MEC-LLM with a trained payload generator to exfiltrate secrets planted in the system prompt, giving AI safety researchers a controlled target for studying internalized covert channels. We release this model organism for future work on monitoring and interpretability of covert communication in language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.