BINGO: Broadening Bias Discovery in LLM Judges via Diversity-Aware Optimization
Abstract
As Large Language Models (LLMs) continue to advance, they are increasingly used as evaluators, making it important to identify biases that influence their judgments beyond response quality. However, existing approaches to bias discovery often remain anchored to known bias patterns, limiting their ability to uncover biases beyond established categories. Moreover, primarily prioritizing strong bias effects can cause the search to repeatedly follow similar successful directions, limiting broader exploration of the search space and leaving other potentially effective bias factors undiscovered. To address these limitations, we introduce BINGO, an agentic framework for broadening automated bias discovery in LLM judges through diversity-aware optimization. Specifically, BINGO employs a dual-loop design: the inner loop promotes semantically and behaviorally diverse candidate bias factors, while the outer loop evaluates how strongly they influence judge preferences. In addition, a reflective feedback mechanism learns from successful, ineffective, and redundant search directions to guide subsequent exploration. Across four LLM judges, BINGO outperforms existing baselines in both bias strength and diversity, improving strength by 19.0% on average over the strongest baseline and reducing semantic and behavioral redundancy by up to 47.0%. Our qualitative analysis further shows that BINGO recovers known biases, refines broader categories, and uncovers novel bias directions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.