Real-Time Streaming Reasoning with Temporal Self-Distillation Policy Optimization
Abstract
Real-time streaming reasoning enables language models to reason incrementally as inputs arrive under progressively revealed observations. While this paradigm reduces response latency, effectively optimizing the intermediate reasoning process remains challenging, as high-quality segment-level supervision is difficult to obtain. Outcome-based reinforcement learning provides a natural alternative because final task success can often be verified directly. However, a streaming trajectory consists of multiple reasoning segments generated under different partial observations, whereas the outcome reward only evaluates the trajectory as a whole. To bridge this gap, we propose Temporal Self-Distilled Policy Optimization (TSDPO), which combines outcome-based policy optimization with on-policy self-distillation. TSDPO leverages training-only privileged information to retrospectively evaluate sampled reasoning segments under the same policy, producing segment-level support signals that are calibrated across streaming positions and used to redistribute trajectory-level advantages over time. This provides fine-grained supervision for streaming reasoning without requiring external teachers or annotated reasoning trajectories. Experiments on Qwen3-1.7B and Qwen3-4B show that TSDPO improves answer accuracy while producing more compact reasoning trajectories, and preserves the low-latency characteristics of real-time streaming inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.