acceptodds
Under review as a conference paper at ICLR 2027

Towards Eliciting Deployment Behaviour in Scheming Models via Steering-Vector Transfer

Abstract

Pre-deployment behavioural evaluations are a critical safety check for LLMs, but scheming models may recognise evaluation contexts and behave differently than in deployment. Some deployments grant consequential affordances that evaluations cannot safely reproduce, a gap we call the safe-to-dangerous shift. These affordances could provide strong evidence of deployment to a scheming model. Prior work proposes white-box steering to elicit deployment behaviour, but constructing a deployment–evaluation steering vector requires candidate-model deployment activations that are unavailable in this setting. Synthetic prompts and historical transcripts provide proxies, yet a scheming candidate may still recognise their use as evaluation and retain its evaluation behaviour. We instead propose constructing steering vectors on already-deployed models, where genuine deployment data exist, and transporting them into the candidate model either unchanged (Direct transfer) or through a linear map fitted by ridge regression on unlabelled text (Ridge transfer). We conduct a feasibility study in a prompted programming organism with a planted deployment-associated behaviour, using three simulated deployed–candidate pairs spanning shared-base, cross-generation, and cross-family transfer. Direct transfer produces passing, typed code in 46.3% of completions on the shared-base pair, close to the oracle's passing-and-typed rate of 46.8%, but reaches only 2.3% on the cross-family pair. Ridge transfer elicits the behaviour on all three pairs, reaching 15.7% on the cross-family pair and approaching the oracle's steering performance on the cross-generation pair (35.7% versus 40.5%). Ridge is stable across seeds on the cross-generation pair, but varies sharply on the other two, performing strongly on some seeds and poorly on others; on the shared-base pair, we trace this instability to the text sample used to fit the map. Steering-vector transfer thus offers an alternative route to eliciting deployment-associated behaviour, complementary to realistic evaluations and candidate-side proxy interventions, but requires further validation in more realistic scheming model organisms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.