acceptodds
Under review as a conference paper at ICLR 2027

Position: Revalidating Probes and Steering Vectors in Scaffold-Evolving Agents

Abstract

Neural interventions primarily measure whether an evaluation changes local agent behavior, but not whether that effect persists after an agent updates, or how the resulting answer is used. This position paper argues that intervention validity must distinguish probe prediction, local intervention effect, and final-answer effect as separate evidence levels. For partial records with valid candidate labels and bounded selector probabilities, we characterize a sharp interval of compatible final-answer effects whose width depends on selector uncertainty and candidate disagreement; indistinguishable observations cannot remove the remaining ambiguity through additional logs of the same kind. For complete records, conditional downstream replay reassesses the effect by executing the updated suffix on each intervention arm while preserving a justified unchanged prefix, subject to preserved upstream laws, a sufficient cut without feedback, and faithful conditional execution. To operationalize this position, we specify a controlled GSM8K study with Qwen2.5-7B-Instruct and judge-only prompt optimization to assess intervention effects and reconstruction, independent-fresh agreement, and total cost. The analysis separates sampling uncertainty from ambiguity requiring new information or renewed computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.