Reinforcing LLM Latent Reasoning with Contextual Loops through On-Policy Flow-Map Self-Distillation
Abstract
Latent reasoning creates a continuous workspace for solving a problem and a context for generating its answer. We conceptually connect recurrent Transformers, continuous diffusion, and finite-time transport and introduce contextual loops, where the same model only recursively refine a latent workspace before performing standard autoregressive decoding. Latent reasoning thereby becomes a source of privileged teacher context whose construction can be controlled through refinement depth and latent space design. We thus introduce joint reward-weighted on-policy self-distillation (OPSD) over the latent and autoregressive states of this process. At student-visited latent states, composed flow maps teach a single update to reproduce several finer refinements. At student-generated texts, the same model uses a more extensively refined latent to supervise its next-token predictions. Verifier rewards weight both objectives, linking the compression of latent computation to the explicit policy that interprets its result. Our instance, LaDi-OPSD, jointly optimizes the two OPSD objectives together with a hierarchical reinforcement objective. The resulting training scheme supports several latent inference loop budgets within one shared model. Our theoretical analysis characterizes reward-weighted supervision and the conditions for transferring teacher quality to a smaller computation budget. Across several coding and mathematical reasoning benchmarks, LaDi-OPSD outperforms other diffusion language models and latent reasoning baselines, while preserving both generation diversity and few-step generation capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.