Embedding Information in Policy Randomness: Distribution-Preserving Implicit Communication in MDPs
Abstract
This paper investigates a new role of randomness in stochastic policies—as a medium for embedding and communicating information. Specifically, we consider an agent interacting with an environment according to a prescribed stochastic policy. By carefully modulating the random action realizations of the policy, information can be embedded into the agent's behavior without altering the prescribed action distribution. We cast this problem as a channel coding problem, in which the controlled system implicitly serves as a communication channel. We then propose a coding scheme, termed posterior-matching randomness modulation (PRAM), that selects actions according to the message to be transmitted. With the aid of common randomness independent of the control system, PRAM preserves the prescribed action distribution for every state and every message. We further prove that PRAM is asymptotically reliable: the message can be decoded with arbitrarily small error probability as the observed state trajectory becomes sufficiently long. Experiments demonstrate the effectiveness of the proposed framework in applications including multi-bit policy watermarking and trajectory provenance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.