acceptodds
Under review as a conference paper at ICLR 2027

DoubleFree: Exploiting Exploration History for Rollout-Free Multi-Task LLMs Training

Abstract

Multi-task post-training can develop domain specialists through independent reinforcement learning (RL) and consolidate their capabilities through multi-teacher on-policy distillation (MOPD). However, standard MOPD generates on-policy student responses and keeps multiple teachers online throughout consolidation, introducing substantial system costs. We view this pipeline as domain-wise exploration followed by cross-domain exploitation. Domain RL produces not only the final specialists, but also the trajectories collected throughout their training. Existing methods primarily use the former while leaving the latter underutilized. We introduce DoubleFree, which reuses these historical trajectories for rollout-free consolidation and precomputes teacher supervision for online-teacher-free student training. To understand why this works, we analyze when updates from historical trajectories can approximate on-policy distillation updates and how combining trajectories across exploration stages can better support an evolving student. DoubleFree requires only minimal changes to existing post-training frameworks. Across math, code, and instruction following on Qwen3-4B, Qwen3-8B, and SmolLM3-3B, DoubleFree achieves - consolidation speedups and reduces GPU-hours by -, while maintaining comparable multi-domain performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.