acceptodds
Under review as a conference paper at ICLR 2027

Beyond Single Traits: Evaluating Multi-concept Steering in Language Models

Abstract

Text generation by Large Language Models can be steered through surface-level prompt engineering (e.g., system prompts) or latent-level representation engineering (e.g., Contrastive Activation Addition). While both can steer individual behaviors, their interaction, compositional limits, and relative influence in multi-concept settings remain poorly understood. This work systematically investigates the scalability, reliability, and adversarial robustness of simultaneous multi-concept steering, focusing explicitly on the interplay between prompt-based and activation-based interventions across language models. Specifically, we study the steerability of LLM's open-ended text generation towards 40 diverse psychological, behavioral, ideological, and personality traits within two experimental setups. First, we evaluate intervention saturation, measuring the operational thresholds at which diverse steering paradigms preserve multi-trait expressions as the number of steering concepts scales. Second, we examine cross-paradigm conflicts by deliberately pitting surface-level system instructions against latent activation steering vectors configured to drive polar-opposite behaviours. We also introduce *hypersteering*, a novel activation-steering paradigm leveraging Vector Symbolic Architectures. At low-to-moderate concept counts, it rivals prompting while generating less cross-trait interference and achieving superior single-trait results. It also surpasses CAA in overriding conflicting system instructions during adversarial attacks. The contributions are both experimental and methodological, from evaluating standard methods in multi-concept open-ended text generation to introducing a novel steering method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.