acceptodds
Under review as a conference paper at ICLR 2027

Thought Latents: Which SAE Features Matter for Reasoning?

Abstract

A reasoning state can lead to different future outcomes, yet a recorded trajectory reveals only one of them. How do internal features reflect these possibilities as reasoning unfolds? We study thought latents: sparse-autoencoder (SAE) features examined through their relationship to future outcomes. We pair boundary activations with repeated continuations and use changes in activation and outcome probability to identify candidates. Our evaluation separates discovery associations, generalization to unseen sources, and decoder-direction interventions. We evaluate these relationships in blackmail and whistleblowing scenarios, including independent-source tests. On held-out whistleblowing histories, two selected directions increase external-reporting decisions by 12.5 and 11.5 percentage points relative to contemporaneous zero, surviving correction over 16 candidates. These measurements connect internal representations to possible futures while distinguishing natural readouts from the effects of applying decoder directions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.