Watch the Model Think: On-Policy Extraction of Activation Steering Vectors
Abstract
When a model solves a problem on one attempt and fails it on the next, what separates the two is rarely the final answer token; it is the trajectory that reached it. Contrastive activation steering leaves that signal unused: CAA, SADI, RepE and ITI build their direction from experimenter-supplied text, recorded while the model reads rather than reasons. That choice also caps what the vector can express, since polarity must be written into the text, and a task judged only by outcome offers nothing to write it with. ROAST makes the trajectory itself the contrast: sample rollouts, let an outcome verifier split them into successes and failures, and contrast the reasoning that worked against the reasoning that did not. The only thing that changes is which text the activations are read from, and a matched teacher-forced control — rollouts, labels, answer text and pair counts held fixed, the trajectory alone stripped — isolates what that buys: on GSM8K at B the pairs alone give points and restoring the trajectories gives , the larger step on all six columns tested and the only one of the two to clear seed variance anywhere. Holding token count and anchor fixed, substituting an equal-length neutral prefix or another question's reasoning falls below no intervention — worse than supplying no reasoning at all. The two corpora are far apart geometrically as well: the two readings of one direction sit a median and apart at the two scales we probe, within of that net of a split-half null. Reading from rollouts forces two corrections — keep the full difference vector, since Top- masking discards half its energy, and give each question one vote rather than one per pair — and grouped aggregation is what keeps GSM8K above baseline under verifier noise. We rest the comparison on the benchmarks scored without an answer parser — GSM8K, MATH500 and IFEval — where ROAST is best in all six cells across two models, by up to over no intervention, for wall-clock and no added context; on the parser-scored multiple-choice benchmarks it also leads every steering baseline on average at all three scales. Across nine benchmarks and nine models (B–B, four families), ROAST improves the per-model average at every scale. https://anonymous.4open.science/r/ORBIT-E014/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.