acceptodds
Under review as a conference paper at ICLR 2027

MupCoder: Mitigating Reward Sparsity in Code Generation via Multi-Perspective Token Level Reinforcement Learning

Abstract

The demands of syntactic correctness and logical precision in code generation remain challenging for off-the-shelf large language models, motivating post-training alignment as a common approach. However, current approaches face critical constraints: Supervised Fine-Tuning (SFT) is prone to in-distribution overfitting and struggles with out-of-distribution generalization, while Reinforcement Learning (RL) suffers from reward sparsity. Therefore, we propose MupCoder (Multi-perspective Coder) and a token-level RL framework that provides dense process supervision without requiring auxiliary reward models. Rather than relying solely on sparse outcome feedback, MupCoder equips the model with multi-perspective hints, such as reference demonstrations and test cases. Conditioned on these contexts, the model acts as its own teacher via on-policy self-distillation, deriving dense token-level rewards from the sampled-token log-probability gap between the standard student branch and hint-conditioned teacher branches. To capture long-horizon dependencies in code generation, we further accumulate these dense signals through a future-aware discounted return, which propagates downstream alignment information to earlier tokens. This dense shaping signal is combined with sparse execution feedback to anchor functional correctness. Experiments on multiple code generation benchmarks show that MupCoder consistently improves generation accuracy over representative baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.