LLM Learning State Engines Improve Reasoning with Self-Supervised Operation Discovery
Abstract
We want LLMs to make explicit logical choices and execute them consistently, providing a basis for stable extrapolation to longer or unfamiliar problems. Checkable reasoning and computation matter for scientific discovery and software security (Romera-Paredes et al., 2024; Klein et al., 2009). Chain-of-thought (CoT) exposes intermediate steps (Wei et al., 2022), but choosing an operation and carrying it out share the same generation process. Rule-based systems obtain explicit control through hand-designed operations. In this paper, we introduce LLM Learning State Engines that learn this interface from transition examples using self-supervised targets. For Qwen3-1.7B, we cluster hidden-state changes between adjacent proof states into anonymous codes and train their use with evidence and next-line targets. The classes align with human proof roles; fixing or permuting their codes eliminates successful proofs. Separate evidence generation and selection support one-line updates. On PrOntoQA (Saparov & He, 2023), the engine improves answer accuracy from 48.33% to 81.67% and exact-process accuracy from 23.33% to 57.78% over Token CoT. Tape experiments test learned execution, new operation sequences, and encoder transfer. With matched Transformer architectures, delegating execution reduces prediction to action selection and improves complete-trajectory accuracy beyond training horizons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.