Cognitive Steering of Language Models
Abstract
Current approaches to identifying representations that influence a language model’s (LM) behavior often rely on noisy datasets or imprecise linguistic descriptions to define complex behavioral profiles. Such approaches can be confounded by spurious features of the data and our own biases in defining their content, resulting in unexpected behaviors at intervention time. In cognitive science, probabilistic generative models of cognition provide formal accounts of variables that parameterize complex human behaviors, including behaviors relevant to LM personality and alignment. In this work, we treat cognitive models as data-generating processes and use their outputs to supervise the search for representations corresponding to their parameters, which we term cognitive variables. We identify these representations in both activation and weight space and use them for cognitive steering: controlling LM behavior according to the computations assumed by the originating cognitive models. According to LM judges, cognitive variables derived from models of interpersonal communication control free generations in naturalistic settings beyond the toy domains the cognitive models were designed to represent, most reliably when distilled into model weights. Our approach provides a principled method for steering naturalistic LM behavior using empirically validated, formal models of human cognition.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.