acceptodds
Under review as a conference paper at ICLR 2027

Showing vs. Telling: Steering with Feedback Vectors

Abstract

Input–output demonstrations can induce activation vectors that execute a task when injected into a language model. Prior work extends this idea to task instructions. We study feedback vectors extracted from rules, explanations, and critiques that specify a mapping without providing input–output examples. Across a suite of seven models spanning four architectures and 1B–9B parameters, we compare textual feedback with three vector constructions, and examine prompt format, representation, and out-of-distribution behavior. Under a Q/A prompt on the primary 1B model, the best of the three vector methods, selected on the training pool, outperforms textual critiques on five of six tasks, and a single fixed method does so on four; across all seven models, that fixed method exceeds critique tokens in 21 of 42 model–task pairs. Feedback-vector accuracy tracks the model's ability to use feedback as text across 146 conditions (). Across three critique prompt templates, however, best-of-methods vector accuracy varies less than accuracy from feedback presented as text. Feedback and demonstration vectors also generalize differently: feedback vectors can outperform function vectors when a regularity in the demonstrations breaks, while function vectors can be stronger on rare factual retrieval. Together with their largely different selected attention heads, these findings are consistent with task representations shaped differently by demonstrations and linguistic feedback.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.