On the Emergence and Effectiveness of User Profiling in Language Models
Abstract
As an increasing portion of humans spend an increasing portion of their time conversing with and directing language models, they will be more subject to the implicit profiling language models are capable of. Understanding how these user profiling capabilities emerge both in the models themselves and during training is essential for privacy and mitigating bias, but existing studies are limited in scope and in realism. We show that language models represent the gender and age of the person writing, that this representation strengthens with scale and forms early in training. Training linear probes on 162k human written blog posts across 29 base models spanning 14M to 32B parameters shows consistent trend lines of about 3 (gender) and 4.4 (age) point increases in accuracy per tenfold increase in size, with a common slope across model families. In OLMo 2, most of the signal appears within the first 3% of pretraining and is mostly unchanged by midtraining and posttraining. Probes trained only on blogs transfer to real user conversations with AI assistants, and transfer improves with model size.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.