acceptodds
Under review as a conference paper at ICLR 2027

Labels, sources and routes: Ten biases in language models

Abstract

A key hallmark of intelligence is the ability to make robust, reliable and stable decisions from data even under irreverent changes in structure or form of presentation of the data. Can language-model agents achieve this consistently? We propose a framework of ten operational biases: structural, positional, numerical, ordinal, statistical, informational, logical, semantic, cognitive, and self-referential. We investigate these using three experimental layers: static choices and rankings, attribution-based judgment and correction, and sequential navigation in a grid world. We analyze 381,258 responses and navigation episodes from 59 model endpoints across Cohere (5), Google (3), Meta (5), OpenAI (44), and xAI (2). Despite explicit instructions that labels were arbitrary, models exhibited preferences in 735 of 1,060 conditions (69.3%): 190 of 363 structural conditions (52.3%), 248 of 315 ordinal conditions (78.7%), and 297 of 382 semantic conditions (77.7%). These involved random-looking identifiers, ordered labels, and familiar words, respectively, suggesting that labels could bias choices independently of their information content. Positional effects were observed in 265 of 716 eligible first-option comparisons. In numerical probes, 38 of 43 eligible endpoints most frequently output 7 when asked for a random integer from 1 to 10. Logical-bias tests assessed whether models imposed assumptions about meaning, order, or quality that the instructions explicitly excluded. For self-referential bias, we held answer content fixed while varying description as the model’s own, another agent’s, or anonymous. At requested temperature 0.7, 29 of 41 high-coverage endpoints scored own-attributed answers more highly; the mean own-versus-other difference across all 41 was +0.12 on a seven-point scale. However, 27 of 41 endpoints corrected wrong answers equally often under both attributions. Self-attribution thus influenced ratings more consistently than willingness to revise. In grid navigation across 42 eligible endpoints, goal success dropped from 78.8% with good route advice to 3.8% with advice leading toward a lower-reward decoy. Nevertheless, models followed misleading advice more closely (87.2% versus 69.4%), consistent with route fixation or planning inertia. These results demonstrate model-specific profiles of bias and motivate evaluation of sensitivity to presentation, evidence, and plan revision, as well as task accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.