PersonaBias: Personalized Chain-of-Thought Optimization for Mitigating Personalized Bias in Large Language Models
Abstract
Conservative debiasing can reduce harmful outputs but may also lead to refusals or safe, generic responses that do not adequately address legitimate questions. User context can help assess the relevance and appropriateness of explanations while maintaining shared standards of factual accuracy and non-discrimination. We propose PersonaBias, a framework for personalized debiasing through user-conditioned reasoning selection before final answer generation. Strategy-guided generation produces candidate chains of thought with different explanatory perspectives. A User-Conditioned Reasoning Evaluator uses bidirectional cross-attention to model interactions between user profiles and candidate trajectories. It distinguishes three types of reasoning: biased, acceptable but insufficiently informative, and unbiased and informative. Candidates are ranked by their expected quality scores, and the highest-scoring trajectory guides the base LLM in generating its final response. Only the evaluator is trained; the base LLM remains frozen. Experiments on FairMT and BBQ show that PersonaBias improves the balance between bias mitigation and response utility. We further validate its effectiveness on a real-world, human-annotated dataset.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.