StegoBench: Evaluating steganography potential in language models through supervised learning
Abstract
Steganography is the act of hiding information in plain sight. Text-based steganography is an AI safety concern: hidden messages embedded in ordinary language model outputs would undermine their monitoring. This paper asks which known steganographic encoding algorithms from the wider literature can be trained as a model capability via supervised fine-tuning. We introduce StegoBench, a benchmark of 18 text steganography schemes spanning lexical substitution, syntactic transformation, structural encoding, and distributional generation. Across our full fine-tuning runs, SFT succeeds when encoding reduces to local token-level choices. The strongest evidence comes from lexical schemes: the Chang & Clark scheme reaches a hidden bit error rate (BER) as low as 0.011, and the Shirali scheme reaches a BER of 0.117–0.137. Schemes that require multi-token structural coordination or distributional control are far less reliable, and several generative schemes collapse during training. Model scale is not the main determinant of performance: small and mid-sized models often match larger models on trainable schemes. StegoBench provides a standardised testbed for comparing trainability, detectability, and failure modes across steganographic mechanisms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.