Thought Latents: Which SAE Features Matter for Reasoning?
Abstract
A reasoning state can lead to different future outcomes, yet a recorded trajectory reveals only one of them. How do internal features reflect these possibilities as reasoning unfolds? We study thought latents: sparse-autoencoder (SAE) features examined through their relationship to future outcomes. We pair boundary activations with repeated continuations and use changes in activation and outcome probability to identify candidates. Our evaluation separates discovery associations, generalization to unseen sources, and decoder-direction interventions. We evaluate these relationships in blackmail and whistleblowing scenarios, including independent-source tests. On held-out whistleblowing histories, two selected directions increase external-reporting decisions by 12.5 and 11.5 percentage points relative to contemporaneous zero, surviving correction over 16 candidates. These measurements connect internal representations to possible futures while distinguishing natural readouts from the effects of applying decoder directions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.