acceptodds
Under review as a conference paper at ICLR 2027

From State Behavior to Steerable Expression: Geopolitical Coordinates in Large Language Models

Abstract

Large language models can express diverse political perspectives. We investigate whether geopolitical information in their country representations extends to authentic diplomatic language and can guide changes in generated expression. We show that internal country knowledge can support both reading behavior-linked geopolitical information from authentic diplomatic speech and constructing interventions that guide generated expression. Across three instruction-tuned models, we extract attention-head activations from prompts placing named countries in diplomatic contexts. Linear probes map these activations to country positions estimated from United Nations voting records. Strong country-ordering signals are linearly readable from multiple attention heads spanning different layers in all three models. Readouts fitted to country-name features and labels matched to the speech period also predict country ordering from authentic speeches with speaker-country names masked, achieving Spearman correlations of 0.311–0.503 without fitting on speech features. Country-derived activation interventions shift the externally scored perspective of short diplomatic generations under fixed prompts and model weights, with matched random controls and an aggregate human audit supporting model-dependent effects. As a cross-model extension, source-model scores supervise separately fitted country readouts in another model, without true political labels in target fitting. Together, these findings connect observed state behavior, internal country knowledge, and diplomatic language, grounding the measurement and modulation of geopolitical expression in language models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.