LARA: Verifier-Guided Rule-Level Semantic Fuzzing for Auditing Local Agent Runtimes
Abstract
Local LLM agents mediate between partially trusted inputs and privileged host actions through runtime code. Historical fixes reveal failures in this mediation, but their code-level forms are project-specific. We present LARA (Local Agent Runtime Auditor), which formulates security-rule synthesis as verifier-guided, rule-level semantic fuzzing: semantic rules replace grammar-generated inputs; applicability and validity reasoning drive residual-conditioned CULL, MODIFY, and AUGMENT mutations rather than fixed operators; and vulnerable/fixed execution feeds new rounds until semantic saturation or the round budget is exhausted. Accepted rules are grounded in a runtime and executed as CodeQL queries. On 275 OpenClaw construction cases, this search raises Patch Recall from 61.1% to 81.5% and patch-discriminating Success from 61.1% to 80.0%; an ablation attributes 49 of the 52 additional successful cases to semantic rule revision rather than query-only refinement. Before target vulnerability feedback, adapted rules localize 17/22 Nanobot and 9/12 Hermes repairs, but discriminate only 5/22 and 4/12, respectively. Refinement on each target’s historical cases raises Success to 15/22 and 7/12. Separately, review of frozen-corpus findings on four latest snapshots produced 25 previously unreported vulnerability candidates validated by proof-of-concept (PoC) exploits. These results suggest that reusable security rules from historical fixes can bootstrap the auditing of new local agent runtimes, for which target-agent-specific evidence and rule adaptation is essential.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.