Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in LLMs
Abstract
Fair decisions require ignoring irrelevant, potentially biasing, information: a decision-maker must approximate what they would have decided had they not known, say, a job candidate's gender or race. This counterfactual self-simulation is notoriously hard for humans, producing biased judgments even by well-meaning actors. We show that large language models (LLMs) share this limitation, in offsetting gender and race biases and in overcoming sycophancy: prompting models to ignore or pretend not to know biasing information fails – and occasionally backfires. Unlike humans, however, LLMs can be given a ground-truth model of their own counterfactual cognition: their own API. Enabling models to query blinded copies of themselves yields fairer decisions and an auditable record of every departure from the blinded verdict.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.