Imagine: Selective Width Expansion for Hard-Token Reasoning
Abstract
Small language models can fail at a few difficult predictions and spoil an otherwise correct answer. Extra reasoning steps can repair these errors, but serial computation makes users wait. We introduce Imagine, a post-training method for selective width expansion within a single Transformer forward. Parallel Slot Expansion gives difficult positions persistent expert states. These states interact across layers and merge only at prediction time. A Lookahead Gain Controller allocates width before the current input enters the model. Dense-to-Hard Training teaches the model to use this capacity, then concentrates it on difficult predictions. Our analysis connects the resulting structured wide network to conditions for lower prediction loss. Across three reasoning benchmarks, Imagine+ 0.6B achieves 5.8% higher average accuracy than plain supervised fine-tuning. Parameter-matched variants also improve average accuracy. Moreover, the single-slot model preserves SFT-level single-user throughput and reduces the mean latency of the slowest 1% of tokens by 34.9% relative to parameter-matched recurrent transformer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.