acceptodds
Under review as a conference paper at ICLR 2027

Real-Time Streaming Reasoning with Temporal Self-Distillation Policy Optimization

Abstract

Real-time streaming reasoning enables language models to reason incrementally as inputs arrive under progressively revealed observations. While this paradigm reduces response latency, effectively optimizing the intermediate reasoning process remains challenging, as high-quality segment-level supervision is difficult to obtain. Outcome-based reinforcement learning provides a natural alternative because final task success can often be verified directly. However, a streaming trajectory consists of multiple reasoning segments generated under different partial observations, whereas the outcome reward only evaluates the trajectory as a whole. To bridge this gap, we propose Temporal Self-Distilled Policy Optimization (TSDPO), which combines outcome-based policy optimization with on-policy self-distillation. TSDPO leverages training-only privileged information to retrospectively evaluate sampled reasoning segments under the same policy, producing segment-level support signals that are calibrated across streaming positions and used to redistribute trajectory-level advantages over time. This provides fine-grained supervision for streaming reasoning without requiring external teachers or annotated reasoning trajectories. Experiments on Qwen3-1.7B and Qwen3-4B show that TSDPO improves answer accuracy while producing more compact reasoning trajectories, and preserves the low-latency characteristics of real-time streaming inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.