acceptodds
Under review as a conference paper at ICLR 2027

Decodable but Confounded: Poker Bluff Labels in Llama-3.1-8B Hidden States

Abstract

A concept label can be highly decodable from hidden states even when the readout largely exploits task variables that make the label predictable. We make this failure mode directly measurable in poker using 44,631 examples annotated by GPT-4o with binary bluff labels. For Llama-3.1-8B-Instruct, we extract pre-response residual-stream activations at the final prompt token and fit layerwise logistic probes with statement-group-disjoint splits and validation-only layer selection. The label is strongly decodable from both an explicit bluff-classification prompt and a question-stripped neutral-state prompt, reaching held-out AUROC 0.936 and 0.935. However, observed poker-state covariates alone reach AUROC 0.920, and removing activation components linearly predictable from those covariates reduces probe AUROC to 0.722. Because poker state is semantically relevant to bluffing, these controls do not isolate “pure bluff information”; they quantify how much of the headline probe signal overlaps directly observed task structure. Probe-direction interventions produce signed, dose-dependent shifts in normalized next-token P(Yes), but at the strongest tested dose no non-degenerate condition changes the model’s argmax Yes/No response, while matched random, shuffled-label, and nuisance directions produce effects of comparable or greater magnitude. A 32,768-feature TopK sparse autoencoder explains 95.2% of held-out activation variance, yet selected single features are weak held-out predictors and learned dictionaries are unstable across seeds. Across probes, interventions, and sparse features, the evidence supports label recoverability and output sensitivity without identifying a bluff-specific mechanism.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.