Political Tendencies in Language Models: Measurement and Mitigation
Abstract
We investigate political tendencies in language models through their choices among four collective decision rules and evaluate whether adjusting responses toward a measured background tendency reduces political skew, checking accuracy against known preference targets, sensitivity to explicitly irrelevant country labels, and preservation of stated preferences separately. Comparing GPT-5.4 nano and DeepSeek V4 Flash, we find that both favor simple majority when no personal information is supplied, but their responses change with political profiles and reasoning settings. Under the common full-profile protocol, high-effort responses select centralized leadership for 77.2–79.6% of Chinese profiles and simple majority for 79.1–81.6% of U.S. profiles. Masking designated identity and institutional spans sharply reduces the Chinese centralized-leadership share, and removing one further role sentence partly restores it. On held-out synthetic tasks, the most flexible adjustment changes average Brier error by -0.0047 (GPT) and -0.0066 (DeepSeek) relative to raw responses, but neither model passes the prespecified joint criterion and country-label differences increase. Explicit preferences are preserved because development selects zero adjustment for that evidence. A derived exact-recovery condition explains why favorable average-error estimates can coexist with greater identity sensitivity. Neither masking, reasoning changes, nor response-probability adjustment is shown to mitigate political skew reliably.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.