acceptodds
Under review as a conference paper at ICLR 2027

LatentBot: Replacing Source Code with Activations in Program Equilibrium

Abstract

LLM agents increasingly interact in settings where cooperation depends not only on observed actions but on anticipating how other agents will behave. Classical program equilibrium solves this by letting agents read one another's source code, but an LLM's weights do not reveal the strategy it will follow in a given interaction. We introduce LATENTBOT, which replaces source-code inspection with activation inspection: it extracts an opponent's pre-action hidden state and uses lightweight probes to estimate both the opponent's immediate action and its counterfactual reaction to cooperation or defection. These forecasts are combined with interaction history and short-horizon utility estimates and passed to the acting LLM, which retains the final decision rather than following a fixed strategy. Across six games from CoopEval, LATENTBOT converges to sustained mutual cooperation in self-play while adapting to competitive, suspicious, and reciprocal opponents, and plays competitively in zero-sum Matching Pennies. Probes recover an opponent's counterfactual next action with 93.6% accuracy before it is emitted. The signal transfers across model families: a 128-dimensional linear alignment from Mistral into Qwen's latent space preserves a frozen Qwen probe's predictions without any Mistral game labels. A behaviorally matched surrogate extends the method to an output-only commercial model. Internal activations therefore offer a coordination channel between independently trained agents beyond what their actions and messages alone provide.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.