RMC-Sec: A Process-Gated Reward Framework for Secure Code Generation
Abstract
Large language models are widely used to assist software development, yet vulnerabilities in generated code remain a major obstacle to their reliable use. Prior work has explored reinforcement learning (RL) to improve code security using feedback from functional tests or static analysis results. However, these outcome-based rewards provide little direct feedback on security reasoning and assign the same reward to responses with the same evaluation outcome, regardless of differences in reasoning quality. To address these limitations, we propose RMC-Sec, a process-gated reward framework for secure code generation. Starting from a policy initialized via supervised fine-tuning, RMC-Sec evaluates three sequential stages—Risk Recognition (R), Mitigation Planning (M), and Implementation Consistency (C)—and integrates the resulting scores into a hierarchical gating reward function. To mitigate reward hacking, we check the final code with a static analyzer and use the binary pass/fail result as a hard gate for reward assignment. We also mix security tasks with functional code generation tasks during RL to balance code security and functional correctness. Experiments show that, with 7,768 RL training examples, the 14B RMC-Sec model achieves the best or joint-best performance across five code generation and vulnerability repair metrics. On BaxBench, it improves the joint functional–security pass rate by 68.0% relative to the strongest baseline, while preserving functional performance on the evaluated general-purpose coding tasks. We further demonstrate that the learned security reasoning generalizes to vulnerability analysis and repair.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.