acceptodds
Under review as a conference paper at ICLR 2027

Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in LLMs

Abstract

Fair decisions require ignoring irrelevant, potentially biasing, information: a decision-maker must approximate what they would have decided had they not known, say, a job candidate's gender or race. This counterfactual self-simulation is notoriously hard for humans, producing biased judgments even by well-meaning actors. We show that large language models (LLMs) share this limitation, in offsetting gender and race biases and in overcoming sycophancy: prompting models to ignore or pretend not to know biasing information fails – and occasionally backfires. Unlike humans, however, LLMs can be given a ground-truth model of their own counterfactual cognition: their own API. Enabling models to query blinded copies of themselves yields fairer decisions and an auditable record of every departure from the blinded verdict.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.