LUA: Functional Alignment of Large Code Models with Streaming Unit Tests
Abstract
Large code models (LCMs) have shown strong performance in solving practical software engineering tasks. However, the generated code may not fully align with the user's specification, motivating further alignment steps through reinforcement learning (RL) based finetuning. Recent RL-based finetuning methods for LCMs exploit code-specific properties, particularly the executable property, by using feedback from unit test executions. Yet, existing methods use a limited set of unit tests. In this paper, we investigate how to effectively leverage a large number of unit tests by rethinking curriculum learning as a unit test selection problem. Whilst prior work has applied curriculum learning to order training problems by difficulty, we utilize it to identify informative unit tests from a large continuously generated pool. To this end, we propose a lightweight test-difficulty learning layer that estimates the informativeness of individual unit tests, together with an adaptive decoding strategy that selects the tests according to the learned informativeness scores. In particular, we study the scaling behavior of unit-test feedback, the efficiency of test selection, and the need for adaptive selection as the informativeness changes during training. We find that simply executing more tests increases computation cost without necessarily improving performance. In contrast, our adaptive selector effectively tracks informative tests from a continuously generated pool, achieving better downstream performance at a lower computational cost. Experiments across diverse LCMs and multiple benchmarks demonstrate that our method effectively leverages large amount of unit tests consistently improves coding performance over standard reinforcement learning based finetuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.