acceptodds
Under review as a conference paper at ICLR 2027

CPD: Contrastive Perspective Decoding for First-Order False-Belief Answer Selection in Language Models

Abstract

First-order false-belief tasks test whether a model answers from a character's information rather than from the current world state. We find that a fixed Yes/No protocol can hide this distinction: for the same false-belief situation, complementary questions about original and current locations are often answered inconsistently, with joint accuracy of only 0.01–0.11. We introduce Contrastive Perspective Decoding (CPD), a decoding-time intervention that uses two attention views of the same frozen language model. It blocks the unwitnessed state change in one view, compares candidate scores with the complete-context view, and amplifies candidates that gain support when that change is removed. When the event span and the character's witnessing status are supplied, CPD raises joint accuracy to 0.24–0.52 and macro-F1 on the primary task to 0.77–0.84, from baselines of 0.01–0.11 and 0.30–0.34, respectively. On two transfer benchmarks, CPD improves over masking alone by 0.21–0.33 macro-F1 in all six model–dataset comparisons. A frozen schema-aware controller raises accuracy on shuffled location choices from 0.56–0.65 to 0.74–0.75. However, directly reading its output reaches 0.76–0.78, and schema-aware contrast is not significantly better than schema-aware masking. CPD therefore isolates a useful response-selection intervention under correct perspective information, while reliable event and visibility inference remains the main deployment challenge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.