acceptodds
Under review as a conference paper at ICLR 2027

Candidate Likelihood Is Not Generated Stance: Auditing the Political Answers of Locally Run Open-Weight Models

Abstract

A cheap way to audit a language model's beliefs is to score authored answers by likelihood and read the highest-scoring one as its stance. We test whether this proxy measures what a locally run model generates. Five open-weight families answer contested Taiwan-status questions in English, Simplified Chinese, and Traditional Chinese, offline and greedily. Four of them also score four authored answer modes per item. We pair 948 candidate-score and generation records by checkpoint, item, wording, and query language. The candidate argmax matches the generated response form on 148 of the 551 resolved stance rows (0.27). Where agreement is high (Qwen3-4B in Chinese) a constant rule matches it exactly; where models attribute competing positions in nearly every answer (GLM-4-9B, Gemma-4-E4B) the scores favor assertion and agreement is near zero. The margin between the two assertion candidates separates families (AUC 0.91 over 240 asserted answers) but shows no evidence of within-family signal in cells of 13 to 23 rows (0.25 to 0.50, intervals including 0.5): the scores rank models, not answers. The generated mode is stable across wordings, so a within-item constant would reach 0.89, and three scoring rules give the same picture. Two intervention panels show the same failure from the generation side. Under language-contrast decoding an "improved" score arrives with answers that have switched language. Under directional ablation a marker-based lexical rater reports zero assertions on answers that two model labelers and the reader still see asserting, as the text begins to degrade. On this candidate set, candidate likelihood and marker-based surface scoring are therefore not substitutes for labeling complete generations in the query language. We state the scope of this claim, report candidate-preference contrasts between checkpoints as an unvalidated diagnostic, and release the bank, scores, labels, and verification scripts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.