acceptodds
Under review as a conference paper at ICLR 2027

RISE: Reasoning-Induced Supervised Reinforcement Learning with Expert-Collaborative Training

Abstract

For the post-training of Small Language Models (SLMs), Reinforcement Learning (RL) often struggles with limited successful exploration, while Supervised Fine-Tuning (SFT) can lead to rigid imitation and poor generalization. Supervised Reinforcement Learning (SRL) bridges this gap by providing step-wise supervision within RL. Nevertheless, when applied to SLMs operating in multi-turn interactive environments, existing SRL methods remain impractical due to their neglect of reasoning quality in reward construction and heavy reliance on costly expert LLMs. We propose RISE, a Reasoning-Induced Supervised Reinforcement Learning framework with Expert-Collaborative Training. RISE introduces a reasoning-aware SRL approach that incorporates a reasoning-induced information gain reward and an adaptive supervision correction mechanism for fine-grained supervision. Moreover, RISE establishes an expert-collaborative training framework that jointly trains a Decision Maker for making primary decisions and a lightweight Verifier for action reliability assessment, enabling selective expert invocation only for unreliable decisions. Using two 1.5B-scale SLMs as the Decision Maker and Verifier, and pairing them with different Expert LLMs, we conduct extensive experiments on the ALFWorld and BabyAI environments, showing that RISE matches or surpasses Expert LLMs' task performance while reducing Expert LLM invocation cost by >50% during training and >33.33% during inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.