acceptodds
Under review as a conference paper at ICLR 2027

Bias Discovery and Steering in VLM Judges

Abstract

Vision-language models (VLMs) serve as multimodal judges in evaluation, training, and inference, but they exhibit biases that propagate to each of these settings. Existing methods for detecting biases typically analyze judge outputs, either testing for specific biases or discovering them through properties of the evaluated responses. Those that do examine the judge's internal states focus on predefined biases. We instead discover biases from these states without choosing them in advance, and correct them by modifying the same states. Using sparse autoencoders, we extract interpretable features from hidden states and introduce a bias score that measures how differently the judge and humans weigh each feature. We then steer the judge along selected features, with a lightweight policy adapting the intervention to each input. Applied to two judge models and two datasets, our method surfaces biases spanning safety, visual grounding, instruction following, and multilingual interaction, including several that are largely unexplored in judges. Steering with bias features produces statistically significant improvements in accuracy for both judges, with increases on relevant task subsets of up to 11.6 percentage points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.