HAIR: Fine-Strand Hidden-State Anchored Implicit Reasoning with Single-Prefill Inference
Abstract
Chain-of-thought (CoT) reasoning is accurate but slow, because every intermediate step is decoded as text. Latent reasoning moves these steps into hidden states, but without dense step-level supervision it stays below CoT beyond 1B parameters, it never commits an intermediate result, and it is compared with CoT baselines that differ in trace format, objective and decoding at once. We present HAIR, a streaming reasoning framework that commits each intermediate result as a one-line textual checkpoint, such as G1=G0*2=18, and places a block of b latent positions before each checkpoint in one append-only KV cache. Goal-boundary distillation (GBD) anchors the hidden state at each checkpoint to a shared-weight explicit path that reads the teacher's rationale for that step. Because the latent blocks and GBD can be removed while the checkpoint stream, data and decoding stay the same, HAIR also tests whether latent computation adds anything beyond the checkpoints. On GSM8K with Qwen3-1.7B, HAIR retains 97.5–99.1% of CoT accuracy at 42–59% of its latency from 18k training records, while CODI and Coconut trained on the same records score 30 or more points lower. The checkpoints carry this accuracy: latent blocks and GBD change it by less than the seed-to-seed spread, although GBD lowers its loss by three orders of magnitude. The savings carry over to SVAMP, GSM-Hard and MATH.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.