acceptodds
Under review as a conference paper at ICLR 2027

Decoupling Ownership Bias from LLM Verification

Abstract

Large language models increasingly rely on verification to evaluate intermediate artifacts and guide autonomous decision making. However, they often judge the same artifact more favorably when it is claimed as their own, creating an ownership bias that can undermine reliable verification. To address this problem, we develop a framework for mechanistic analysis and inference-time intervention on ownership bias. We first isolate the effect of claimed authorship by holding the evaluated artifact fixed while varying only its attribution. We then identify an ownership-associated direction from paired internal activations and project it out during inference, while comparing against correctness-related controls. Across multiple models and tasks, this single-direction intervention substantially narrows the ownership gap while largely preserving correctness discrimination, demonstrating effective generalization to new benchmarks and attribution phrasings. These findings reveal a selectively removable internal component of ownership bias, providing a mechanistic basis for more impartial and reliable LLM verification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.