acceptodds
Under review as a conference paper at ICLR 2027

NEST: Training Continuous Latent Reasoning as Anytime Computation

Abstract

Adaptive latent reasoning aims to allocate computation according to input difficulty, but early stopping is only useful when intermediate latent states are themselves reliable answer endpoints. We study this mismatch and propose NEST (Nested Endpoint Supervision for Trajectories), an anytime-training framework that trains variable-depth competence before learning variable-depth execution. NEST supervises multiple prefixes of a single shared recurrent latent trajectory to produce complete answers, so each intermediate state serves both as a valid prediction endpoint and as the continuation state for deeper reasoning. The full method further uses decision distillation, ordered alignment, and monotonic refinement to structure this hierarchy, followed by a calibrated student-only policy for adaptive execution. Across eight reasoning benchmarks with Qwen3.5-4B, NEST achieves 82.68 average accuracy at 0.402 relative GPU cost, compared with 82.07 accuracy at 0.592 cost for SwiReasoning. At matched accuracy, NEST reduces GPU compute by 38.3%; at matched GPU cost, it improves accuracy by 0.95 points. Fixed-prefix and matched-router controls show that the gains arise from improved shallow-prefix competence rather than stronger halting machinery. Adding Core NEST to SIM-CoT and CODI similarly improves accuracy while reducing latent depth, and pool-free deployment with disjoint calibration preserves the operating regime. These results support nested endpoint supervision as an effective strategy for adaptive continuous latent reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.