Bias Discovery and Steering in VLM Judges
Abstract
Vision-language models (VLMs) serve as multimodal judges in evaluation, training, and inference, but they exhibit biases that propagate to each of these settings. Existing methods for detecting biases typically analyze judge outputs, either testing for specific biases or discovering them through properties of the evaluated responses. Those that do examine the judge's internal states focus on predefined biases. We instead discover biases from these states without choosing them in advance, and correct them by modifying the same states. Using sparse autoencoders, we extract interpretable features from hidden states and introduce a bias score that measures how differently the judge and humans weigh each feature. We then steer the judge along selected features, with a lightweight policy adapting the intervention to each input. Applied to two judge models and two datasets, our method surfaces biases spanning safety, visual grounding, instruction following, and multilingual interaction, including several that are largely unexplored in judges. Steering with bias features produces statistically significant improvements in accuracy for both judges, with increases on relevant task subsets of up to 11.6 percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.