acceptodds
Under review as a conference paper at ICLR 2027

Steering Language Models Without Contrastive Pairs via Unsupervised Latent Region Matching

Abstract

Activation steering methods control language model behavior by intervening on internal activations, but many steering-vector approaches require contrastive prompts or paired good and bad completions. We study a lower-supervision setting where only a behavior dataset and the model's own zero-shot choices are available. We introduce Unsupervised Latent Region Matching (ULRM), which splits examples into successful and unsuccessful subsets and learns a nonlinear map from unsuccessful prompt-side decision activations toward the successful activation region. Our loss objective encourages the intervention to shift activations toward preferred behavior while preserving task-relevant information. We evaluate ULRM on two sycophancy datasets using Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct. We compare two vector baselines. MeanDiff subtracts the mean unsuccessful activation from the mean successful activation in the same split. Contrastive Activation Addition (CAA) averages paired preferred-minus-nonpreferred completion activations. Despite CAA's stronger paired supervision, our best-performing method, MMD-ULRM, leads both baselines in all eight comparisons, averaging 0.604/0.541 accuracy/probability versus MeanDiff's 0.527/0.487 and CAA's 0.503/0.457. These results suggest that behavior-split latent-region matching is a promising route to practical activation steering without paired contrastive completions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.