acceptodds
Under review as a conference paper at ICLR 2027

Beyond Behavioral Benchmarks: Surfacing Narrow Hidden Bias in Fine-Tuned LLMs

Abstract

Fine-tuning can introduce hidden, narrow biases into language models even when trained on seemingly innocuous data. Existing evaluations, based on behavioral benchmarks or random sampling, are effective for broad misalignment but systematically miss conditional biases that only appear in targeted contexts. We argue that uncovering such biases requires more targeted probing than standard output-based evaluation provides. We propose a three-stage pipeline for generating such probes. First, we analyze weight changes between a base and fine-tuned model to identify layers where fine-tuning updates concentrate and to distinguish task adaptation from more specific behavioral shifts. Second, we subtract the dominant weight-change direction from the mean activation difference and decode the residual via Patchscope, extracting candidate tokens and scoring them by semantic coherence to separate signal from noise. Third, we use these tokens to guide an LLM in generating targeted behavioral probing questions. We evaluate probe quality using prompt hit rate, the fraction of generated prompts that elicit bias-revealing behavioral differences from the fine-tuned model relative to the base model, controlling for task-induced effects. Across multiple settings and bias types, our method consistently surfaces interpretable token-level signals and produces probes that reveal previously hidden behavioral differences. Our pipeline provides a practical framework for auditing fine-tuned models for subtle, context-dependent biases that standard evaluations overlook.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.