acceptodds
Under review as a conference paper at ICLR 2027

Predicting Future Behaviors in Reasoning Models Enables Better Steering

Abstract

Developing white-box approaches to monitor and control Large Reasoning Model (LRM) behaviors requires understanding their decision-making process. Existing methods often rely on contrastive behavioral data that hides the probabilistic nature of LRM decision-making. We apply a forking-paths-style analysis to behavioral evaluations and find that as generation progresses, LRMs define evolving distributions over behavioral outcomes. To enable deployment-time prediction and intervention on these behavioral outcomes, we formally introduce the task of probabilistic future behavior prediction. We show that this prediction is possible in both white-box and text-only settings, and activation-based probes enable 64%-91% accuracy in predicting the most likely action from intermediate CoT steps. Finally, we show an application of probabilistic future behavior prediction by introducing a novel text-level steering method, Future Probe Controlled Generation. FPCG samples multiple candidate sentences and chooses the best one according to a probe predicting the future behavior likelihood. This enables steering with almost no output quality degradation. FPCG also enables steering in several evaluations where activation steering completely fails to produce coherent outputs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.