Do Language Models Model You? Estimating Implicit Beliefs and Their Generative Effect
Abstract
Language models often infer user traits, such as gender, age, or wealth, from conversational cues. These implicit assumptions can silently shape responses, leaving users unaware that any profiling occurred. Auditing this behavior requires answering two tightly coupled questions: *what did the model believe?* (**Q1**) and *did that belief shape the output?* (**Q2**). To answer **Q1**, we evaluate five methods for belief readout across activations, logits, and text. We find that as model size increases, probes trained on internal activations perform no better than eliciting the belief directly from the model's output. To address **Q2**, we adopt a causal framework. We edit either the prompt or the model's activations to reach a target belief, and measure how the answer changes. Using the Gumbel-Max trick, the intervened model reuses the sampling noise of the original answer, so the change reflects only the intervention. The **Q1** readout serves as a "thermometer" that verifies whether each intervention reaches the target belief without changing other beliefs. We find that input-level interventions are more selective than activation steering at changing only the target belief. Finally, we use this framework to audit which beliefs 8 models infer from a user's occupation alone (60 jobs), and whether these beliefs change their answers. We find that job titles alone dictate strong monocultural stereotypes about gender and wealth. Crucially, these inferred beliefs can silently alter the substance or the style of the answers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.