acceptodds
Under review as a conference paper at ICLR 2027

MetaVul: Learning Transferable Agent Harnesses for Vulnerability Detection

Abstract

Large Language Model (LLM) agents are increasingly used to automate vulnerability detection, but reliable detection remains a challenging problem because agents must distinguish subtle security-relevant differences from benign code variation. While underlying LLM capability has improved significantly, agent accuracy also depends heavily on the harness: the code that defines the scaffolding around the LLM. Harnesses are typically fixed and largely hand-written, limiting adaptation to evaluator feedback. Additionally, harness improvements for one dataset or model may fail to transfer to others. To address these challenges, we introduce MetaVul, a self-improving framework for harness discovery in vulnerability detection. Starting from a seed structured reasoning scaffold, MetaVul uses a workflow-based optimization loop to iteratively refine the harness using feedback from paired vulnerable and patched functions, all while preserving the structured reasoning and verdict format. At each iteration, it searches over the full history of prior candidate harnesses, with access to their trajectories, logs, and metrics, and it returns a set of Pareto-optimal harnesses rather than a single best candidate. We apply MetaVul to discover harnesses for architectures including Codex CLI and Claude Code. A validation set (202 pairs) from TitanVul and CWE-Eval is used for optimization, while the unseen PrimeVul dataset (430 pairs) is kept as the held-out set. We later test whether harnesses learned from one model transfer to others not seen during optimization. Results on both frontier and small LLMs show that the optimized harnesses improve pairwise vulnerability detection up to 23.3 percentage points on the validation set (Opus 4.6) and up to 25.8 points on the held-out test set (DeepSeek V4 Flash). Harnesses learned with one model also transfer to others. We show that harnesses optimized with Opus 4.6 transfer to Opus 5, improving held-out pairwise accuracy by 14.7 points without the need for further optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.