acceptodds
Under review as a conference paper at ICLR 2027

Where Does Bias Enter RAG? A Causal Attribution Framework

Abstract

Large language models (LLMs) are increasingly used to support or automate consequential decisions, and such systems often leverage historical data through retrieval-augmented generation (RAG). Given a query case, similar past cases are retrieved and an LLM reasons over them before producing a decision. When the final decision exhibits a disparity between demographic groups, a fundamental question arises: is the disparity inherited from the historical data, induced by the retrieval algorithm, or introduced by the LLM's reasoning over the retrieved evidence? Existing bias audits measure the disparity of the deployed system as a whole, and are therefore unable to attribute it to individual components. In this paper, we make two contributions towards answering this question. First, we develop an attribution framework that decomposes the disparity of a RAG pipeline into the contributions of (a) the retrieval database, (b) the retrieval algorithm, and (c) the LLM aggregator, by modeling the pipeline as a sequence of predictors related to the same underlying causal model. This attribution method applies to any disparity measure defined on the score itself. Second, for causal fairness measures, we derive decomposition results in which each component's contribution is further attributed to the direct, indirect, and spurious causal pathways, yielding a granular bias analysis of the pipeline. We establish identification conditions for the required quantities, and construct estimators that accommodate free-text mediators such as clinical notes or loan descriptions. We demonstrate the framework in two real-world settings, emergency department triage and microfinance loan applications.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.