Towards Unsupervised Removal of Biased Reasoning in LLM and Humans Alike
Abstract
Rational agents should change their beliefs in response to new information. Still, the martingale property of Bayesian beliefs says that the direction of these changes should not be predictable from their initial beliefs alone. We use the Martingale Score (MS), which measures this predictability, to introduce Martingale Training (MT), an unsupervised training signal, and find that it shows promise. On a forecasting model deliberately distilled with confirmation bias, unsupervised MT improves the held-out Brier score from 0.291 to 0.234 on Qwen3-32B and 0.364 to 0.257 on Llama.3.1-8B roughly matching supervised interventions (direct Brier-score training) which took Brier score from 0.364 to 0.246 on Llama 3.1-8B. In simulated human–AI conversations, MS detects when a sycophantic chatbot entrenches its users, an effect an order of magnitude larger than any entrenchment in the chatbot's own beliefs. Training the chatbot on its user’s MS improves the user’s forecasts and reduces entrenchment, but not better than a simple prompt baseline. As language models increasingly shape what people believe, methods that detect and reduce entrenchment could help build assistants that improve reasoning and human rationality, rather than reinforce existing views and beliefs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.