acceptodds
Under review as a conference paper at ICLR 2027

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has achieved great success in developing Large Language Models (LLMs) with chain-of-thought rollouts. It is used for many tasks such as math and coding. Nevertheless, RLVR struggles with sample efficiency on difficult problems where correct rollouts are hard to generate. Prior works propose to address this issue via demonstration-guided RLVR, i.e., to conduct Supervised FineTuning (SFT) when RL fails. However, SFT often requires a lot of data, which can be expensive to acquire. In this paper, we propose FEST, a FEw-ShoT demonstration-guided RLVR algorithm. It attains compelling results with only demonstrations randomly selected from an SFT dataset. We find that three components are vital for the success: supervised signal, on-policy signal, and decaying weights on the few-shot SFT dataset to prevent overfitting from multiple-epoch training. Across several benchmarks, FEST outperforms baselines in the few-shot setting and even matches the performance of baselines trained with orders of magnitude more SFT data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.