MoldCoder: Automated Construction of Code Instruction Data Using Real Test Execution as Functional Supervision
Abstract
High-quality code instruction-tuning data are an important foundation for improving the code generation capabilities of pretrained language models. Beyond data scale, task complexity, and data diversity, code instruction-tuning data also require that the code correctly implement the functional requirements specified by its corresponding instruction. Many execution-based data construction methods use model-generated tests to verify synthetic code, with execution outcomes guiding sample filtering or code repair. However, the reliability of this functional supervision remains constrained by the quality of the generated tests. A central question is therefore how to obtain more reliable functional supervision and leverage it more effectively. To this end, we introduce Mold-Instruct, an automated code instruction data construction method that uses real-test execution as functional supervision: it first recovers repository-dependent real tests into standalone executable test–code pairs, then uses functional requirements from the tests to guide instruction construction and execution outcomes for verification and heuristic capability diagnosis, thereby constructing target-specific training data. By adjusting its extensible configurations, Mold-Instruct constructs 46.7K high-quality test–code pairs from open-source code repositories, which are then used to construct training data for two base models and fine-tune the corresponding MoldCoder models. MoldCoder improves all nine evaluated metrics on Qwen3-4B-Base and all seven on SmolLM3-3B-Base, increasing their respective average scores from 45.7 to 51.6 and from 33.1 to 36.9. Further controlled experiments show that real tests outperform target-model-generated tests in our setting, functional constraints and execution verification each provide gains, and capability diagnosis guides supervision selection that further improves performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.