acceptodds
Under review as a conference paper at ICLR 2027

Causal Intervention on Latent Confounding Bias Neurons in Large Language Models

Abstract

Pre-training corpora of large language models (LLMs) often contain latent confounding information, which can be encoded into model parameters and induce confounder-related neuronal activations during inference. Existing causal debiasing methods for LLMs predominantly rely on the front-door adjustment criterion, which assumes that latent confounders do not directly affect the reasoning process. In practice, however, latent confounders may also affect the reasoning process, violating the assumptions required for valid front-door adjustment. To address this limitation, we propose CILC (Causal Intervention on Latent Confounding Bias Neurons), a neuron-level causal intervention framework for mitigating latent confounding bias in LLMs. Specifically, CILC first employs a Variational Autoencoder (VAE) to infer latent representations from the data and maps these representations into the LLM’s input embedding space via a latent-to-soft-prompt transformation. It then constructs contrastive inputs with and without the resulting latent soft prompts and applies Backward Integrated Gradients (BIG) to localize neurons that are particularly sensitive to latent confounding information. Finally, CILC performs targeted causal interventions on the identified neurons to attenuate the influence of latent confounding on model predictions. Extensive experiments on three LLM backbones and eight real-world datasets demonstrate that CILC effectively mitigates latent confounding bias while consistently improving model performance. Our source code is available at https://anonymous.4open.science/r/CILC-3E63.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.