acceptodds
Under review as a conference paper at ICLR 2027

COMPASS: Navigating Prompt Latent Space for Bias Estimation in Large Language Models

Abstract

Addressing bias in Large Language Models (LLMs) is critical for their reliable and responsible deployment. Current evaluation frameworks primarily rely on predefined benchmarks or fixed sets of prompts to elicit and quantify biased model behavior. However, these approaches cover only a limited set of interactions, leaving largely unexplored how bias varies across the broader prompt space. Characterizing bias beyond such predefined sets is challenging, as the prompt space is effectively unbounded and exhaustive probing is computationally infeasible. To address this challenge, we introduce COMPASS, a framework that formulates bias assessment as an active exploration problem over a learned continuous latent representation of the prompt space, whose geometry is explicitly structured to preserve semantic consistency and bias closeness. Over this learned space, bias and uncertainty are estimated through a probabilistic surrogate model and used to guide the exploration toward regions of either high predicted bias or high uncertainty. The resulting iterative process enables bias estimation beyond the explicitly evaluated prompts, replacing exhaustive exploration with selective LLM probing. Experiments on synthetic and empirically estimated biases show that the learned latent geometry captures both semantic and bias structure across different configurations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.