acceptodds
Under review as a conference paper at ICLR 2027

Reasoning Under Intervention: Learned Soft Prefixes as Behavioral Probes

Abstract

Correct answers alone do not reveal how a language model organizes reasoning. We study this question with learned soft prefixes, short continuous vectors optimized while the model and the reasoning problem remain fixed. Rather than decoding the prefix itself, we interpret it through its behavioral effects by asking which inputs it changes, which it leaves unchanged, and whether those effects depend on meaningful properties of the problem. We apply this framework to syllogistic reasoning in Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B. Broad-target prefixes show that all three models can be steered, while also revealing different sensitivity to the answer interface. The Qwen models remain strongly steerable under randomized answer mappings, whereas Gemma depends more strongly on how the answers are represented. Broad steering alone, however, does not establish that an intervention responds selectively to the logical problem. We therefore pair valid syllogisms that share the same correct answer but differ in whether they rely on the nonempty-term assumption. In both Qwen models, learned prefixes separate these groups strongly, while Gemma shows little comparable effect. Their answer margins also shift differently, which a uniform answer-score bias cannot explain. A series of controls shows that the effect is stronger than for random groupings, is not tied to one presentation, and changes with the stated logical convention, while failing to transfer to an equivalent task formulation. These results give the three models distinct behavioral profiles in syllogistic reasoning. Together, this shows that behavioral intervention probes can reveal differences in how models organize reasoning that are not visible from traditional benchmarks alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.