acceptodds
Under review as a conference paper at ICLR 2027

RL Unlocks Compositional Reasoning via Soft Tokens

Abstract

Large language models (LLMs) acquire reasoning capabilities in post-training via methods such as supervised fine-tuning (SFT) on demonstrations and reinforcement learning (RL) with verifiable rewards, but continuously extending these capabilities without forgetting previously acquired ones remains challenging. Skill neologisms (SNs) have recently been proposed as a way to learn new skills in modular fashion via semantically integrated soft tokens. However, prior work has demonstrated SNs only for direct input–output skills, whereas soft token methods are known to struggle with multi-step reasoning. It therefore remains unclear whether SNs can encode reasoning procedures that compose with other capabilities as effectively as skills internalized in model weights. We study this question by controlling which reasoning skills are available to a base model, training new skills via SNs and weight-based fine-tuning (FT), and evaluating their composition abilities with in- and out-of-distribution skills. Beyond preventing forgetting of existing skills by construction, we find that the effectiveness of reasoning SNs depends strongly on both model capability and training method. Under SFT, more capable models increasingly close the gap to internalized skills on in-distribution compositions, but remain weak on out-of-distribution composition. Training SNs with RL significantly improves performance, matching or exceeding internalized skills on in- and out-of-distribution compositions, and unlocking composition between independently learned SNs. We extend our findings to a medical reasoning task, where RL-trained neologisms significantly improve transfer to OOD compositions over SFT. Our results show that RL-trained skill neologisms are a promising path for continually extending the reasoning capabilities of strong LLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.