Construct Before Compose: Preserving Conditional Structure in Representation-Based Bias Mitigation
Abstract
Social biases in large language models can vary across combinations of demographic attributes, yet existing representation-based interventions often construct directions from isolated attributes and attempt to recover such variation only during composition. However, marginal direction construction can discard cross-attribute information before intervention is applied. Hence, we propose Construct Before Compose (CBC), a condition-aware framework that constructs attribute directions while preserving cross-attribute conditions and then applies composition mechanisms. We evaluate CBC on race–gender, age–gender, and age–race settings across four instruction-tuned models from three model families. Compared with single-attribute construction under the same composition rule, CBC improves the mean conditional bias reduction from 7.70% to 12.13% across the three attribute pairs. Further analyses show that mitigation on marginal attributes does not reliably transfer to conditional targets, and that conditional evaluation is necessary to reveal residual slice-level disparities and cross-attribute interaction effects. Across six general-language benchmarks, CBC maintains comparable utility, with a four-model macro-average accuracy change of -0.30 percentage points relative to no intervention. These results establish condition-aware direction construction as a fundamental design principle for reliable representation-based intervention in large language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.