Computational Reasoning Machine
Abstract
Contemporary work demonstrates that reasoning strategies can structure transformer inference, showing improved task-performance over pure streams of tokens. These frameworks enable : The reasoning processes feature causal dependency structures, procedural abstractions, and task concurrency. However, prior work implementing reasoning strategies requires ad-hoc programmatic wrappers, inflexible and blind to the problem at hand. All of these concepts have analogies to computational primitives in computer systems (e.g. programming languages, operating systems, and computer architecture). We demonstrate that, instead of requiring a model to reason solely in natural language in a stream of tokens, we can augment the model's vocabulary with computational primitives: variables with scopes, functions, and deferred execution. Doing so enables a model to manage its context (reducing the number of tokens attended to), to abstract and reuse its own reasoning process (reducing the number of generated tokens), and to multitask (increasing the decoding throughput) natively during the course of its decoding process. We augment the transformer architecture to produce what we term a to elicit computational reasoning naturally in transformers. Specifically, we modify self-attention to reflect the machine's causal dependency structures, reuse procedural abstractions, reveal task concurrency opportunities, and reject invalid executions of the machine. Our results show a reduction in mean attended tokens per token (39.0-60.8%) and an increase in available decoding parallelism (1.1-1.28x), at the cost of increased KV cache usage (2.6-40%) and a 1.9-17.7pp drop in accuracy. RL narrows the accuracy drop to 1.2-6.7pp.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.