acceptodds
Under review as a conference paper at ICLR 2027

BWB-KD: Domain Knowledge Distillation via Knowledge Injection, Reasoning Organization, and Activation

Abstract

Transferring domain expertise from black-box teachers while preserving general capabilities remains a challenge in domain adaptation. Direct supervised adaptation can improve domain performance at the cost of broader capabilities, while subsequent reinforcement learning does not reliably resolve this trade-off. We propose BWB-KD, a stage-coordinated distillation framework that connects knowledge injection, reasoning organization, and outcome activation through complementary reasoning modes. Non-thinking supervised fine-tuning first builds a domain expert from black-box demonstrations. Thinking-mode on-policy distillation then uses the frozen expert to guide a base-initialized student along student-generated reasoning trajectories. Finally, non-thinking reinforcement learning optimizes the transferred student using outcome feedback. This off–on–off design places expert-guided reasoning organization between knowledge acquisition and task optimization. Experiments on Qwen3.6-27B across three biomedical tasks show improved domain performance over supervised fine-tuning followed by reinforcement learning, with larger gains in non-thinking inference. General-capability evaluations further show better non-thinking retention. Comparisons with offline distillation, alternative reinforcement learning methods, and different stage orders and reasoning modes support the value of coordinating how domain knowledge is acquired, transferred, and reinforced.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.