SABER: From Vulnerability Reports to Executable Repair Tasks at Scale
Abstract
Training reliable vulnerability-repair agents benefits from executable tasks that verify both security and functionality. However, existing resources remain limited in scale or executability. We introduce (calable gent-ased xecutable epair), a pipeline that turns public vulnerability evidence into repository-level repair tasks, each with a reproducible environment, an upstream fix, and behavioral tests. In particular, first links NVD and GHSA records into deduplicated targets, whose complementary fields fill gaps left by either source alone. A tool-using agent then interprets each target's irregular evidence and synthesizes a candidate repair task. It can adapt to heterogeneous project contexts without project-specific rules or manual reconstruction. Quality control further involves deterministic processing, layered validation, and feedback-driven refinement. Applied to 60,000 targets, promotes 22,089 tasks across 10,163 upstream repositories and more than 60 languages. Sampled ablations demonstrate the benefits of combined evidence, complementary validation gates, and two-layer trajectory filtering. Fine-tuning Qwen3.5-4B and Qwen3.5-35B-A3B on filtered trajectories improves vulnerability-repair performance across multiple programming languages and evaluation criteria.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.