Random Choice in Language Models: A Linear Choice State and Low-Rank Sampling Control
Abstract
Large language models (LLMs) often make highly biased choices when asked to select an item at random. We ask whether this behavior is only an output-level failure or whether models internally represent random choice as a distinct task mode that can be read and controlled. We study this question across twelve answer domains using linear probes, sparse autoencoders (SAEs), and causal interventions on internal activations. We identify a domain-general linear representation that separates random-choice requests from deterministic ones, generalizes to unseen domains and paraphrases that do not explicitly mention randomness, extends to broader arbitrary-choice contexts, and replicates across Gemma-2, Gemma-3, Llama-3.1, Qwen2.5, and Qwen3. Causal interventions on Gemma-2-9B, with a qualitative replication on Llama-3.1-8B, reveal a functional decomposition: one direction acts as a switch into the random-choice mode, whereas a separate direction shapes the resulting answer distribution. This decomposition enables a mechanism-guided intervention that improves sampling without task-specific gradient training. We further find that a trained low-rank adaptation (LoRA) policy produces low-dimensional activation changes, motivating a low-rank representation fine-tuning (LoReFT) intervention that often reaches the trained adapter's calibration regime with substantially fewer parameters, though with lower robustness on some skewed targets. Together, these results show that random-choice behavior is supported by compact, causally relevant internal structure that can be exploited for efficient sampling control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.