Learning to Generate Efficient Code with Privileged Efficiency Context and AST-Guided Supervision
Abstract
Large language models (LLMs) have achieved remarkable success in code generation, yet functionally correct programs often present sub-optimal inefficiency in runtime efficiency and memory usage. Recently, there are some method have attempted to alleviate this problem. While effective, these methods only provide sparse feedback at the program level after generation is complete, making it difficult to provide fine-grained supervision during generation. To address this limitation, we propose EffiCode-OPD, an on-policy distillation (OPD)-based framework for efficient code generation to provide efficiency-aware token-level supervision during training. Specifically, EffiCode-OPD samples on-policy rollouts from the student model and uses a stronger teacher model to provide token-level supervision conditioned on an efficiency context that includes reference code, runtime, memory usage, and AST information. In addition, we design an AST-guided token weighting scheme that focuses the distillation objective on efficiency-relevant code segments. Experiments on representative code efficiency benchmarks show the effectiveness of EffCode-OPD for generating efficient code. For example, EffiCode-OPD reduces average runtime by up to 17.9% and memory usage by up to 19.4% relative to the strongest baseline on EFFIBENCH-X, while preserving functional correctness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.