Steering LLM Coding Behaviors with Concept Vectors
Abstract
Large language models follow coding-practice instructions, but prompt-only control neither reveals internal structure nor avoids repeating requirements. We test whether coding behaviors admit useful residual directions. Using factorized multilingual prompts, we generate and teacher-force code from six Qwen2.5-Coder-Instruct models (0.5B–32B) and estimate compliance-conditioned post-token mean differences for commenting, logging, detected defensive constructs, and programming language. The prompt-conditioned contrasts show coherent lexical and cross-language structure. Adding them after the first sampled token changes held-out outputs, but effects depend on model and layer, strong scaling harms validity and correctness, and axis specificity remains unresolved without matched-norm causal controls. Among endpoint-compliant tasks, explicit comment prompts shift outputs along the training-derived direction. Pre-summed vectors share one hook and one residual addition per steered decode call. In one correctness-screened 7B case, prompting and steering both reach 100% detected compliance; steering uses 17.6% fewer input but 5.1% more output tokens (9.0% fewer total), with 0.43 Python pass versus 0.44 for prompting. None of six tested positive strengths qualifies for a logging-plus-defensive profile. These empirical controls require model- and layer-specific calibration and executable evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.