Steering Language Models Without Contrastive Pairs via Unsupervised Latent Region Matching
Abstract
Activation steering methods control language model behavior by intervening on internal activations, but many steering-vector approaches require contrastive prompts or paired good and bad completions. We study a lower-supervision setting where only a behavior dataset and the model's own zero-shot choices are available. We introduce Unsupervised Latent Region Matching (ULRM), which splits examples into successful and unsuccessful subsets and learns a nonlinear map from unsuccessful prompt-side decision activations toward the successful activation region. Our loss objective encourages the intervention to shift activations toward preferred behavior while preserving task-relevant information. We evaluate ULRM on two sycophancy datasets using Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct. We compare two vector baselines. MeanDiff subtracts the mean unsuccessful activation from the mean successful activation in the same split. Contrastive Activation Addition (CAA) averages paired preferred-minus-nonpreferred completion activations. Despite CAA's stronger paired supervision, our best-performing method, MMD-ULRM, leads both baselines in all eight comparisons, averaging 0.604/0.541 accuracy/probability versus MeanDiff's 0.527/0.487 and CAA's 0.503/0.457. These results suggest that behavior-split latent-region matching is a promising route to practical activation steering without paired contrastive completions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.