Promises and Pitfalls of Multi-Turn Refinement for Secure Code Generation
Abstract
AI coding assistants are now widely adopted, yet they frequently generate vulnerable source code. A promising approach to mitigate these issues is inference-time code refinement, where models iteratively improve outputs using feedback signals such as test cases, static analysis, or natural language critiques. However, the effectiveness of different feedback modalities and the fundamental limitations of this paradigm remain poorly understood. We present a comprehensive empirical study of multi-turn refinement for secure code generation across two axes: generator capability and feedback signal. Our experiments across four feedback types and two benchmarks show that iterative refinement improves secure code generation performance, but no single feedback signal uniformly dominates. We identify reward hacking as a primary bottleneck limiting progress where language models often exploit weaknesses in feedback signals by generating superficial fixes that satisfy the critic without addressing the underlying correctness or security issues. We present a taxonomy of reward hacking behaviors and show that these behaviors can become more pronounced with additional refinement iterations. These findings highlight a key trade-off in feedback-based code refinement: while feedback can improve performance, it can also incentivize undesirable optimization strategies. Addressing reward hacking is therefore essential for advancing reliable and secure AI-assisted code generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.