RLPBE: Reinforcement Learning with Adaptive Multi-Aspect Rewards for Programming-by-Example
Abstract
Programming-by-Example (PBE) enables the automatic synthesis of programs from input–output examples, allowing users to implement their intended programs without requiring programming expertise. Classic PBE methods mainly depend on domain-specific languages (DSL) and search strategies to generate programs, but are limited by DSL expressiveness. Recent approaches based on large language models (LLMs) reduce this dependency and synthesize programs from examples through prompting or supervised fine-tuning (SFT). However, the lack of fine-grained supervision signals often leads LLMs to generate programs that deviate from intended behaviors or contain redundant code. To address these limitations, we introduce RLPBE, a reinforcement learning framework designed for PBE tasks. RLPBE incorporates multi-aspect reward metrics covering four aspects: function usage, code summarization, pass rates, and code conciseness. These rewards are aggregated using our proposed performance-guided adaptive reward strategy and utilized to train LLMs through Group Relative Policy Optimization (GRPO). Experimental results on representative PBE tasks demonstrate the effectiveness of our method. For example, on the List transformation task, RLPBE outperforms the strongest baseline by up to 6.43% on Pass@1 and 7.45% on Pass@5.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.