Beyond the Diff: Cost-Aware Agentic Traversal for Industrial Grade Security Review
Abstract
When state-of-the-art LLMs review code for security vulnerabilities, the true cost bottleneck is not finding bugs: it is adjudicating them to ensure minimal false positives. In production scanning, candidates generated by the scanners - probabilistic as well as deterministic - trigger multi-stage deliberation cascades (cross-vendor debate, rebuttal, and arbitration) where adjudicators read exhaustive repository context to prevent developer fatigue from false alarms. Strikingly, telemetry reveals that judges read 3.9× defects trigger multi-stage deliberation cascades (cross-vendor debate, rebuttal, and arbitration) where adjudicators read exhaustive repository context to prevent developer fatigue from false alarms. Strikingly, telemetry reveals that judges read 3.9× more changed files than candidate findings cite, burning over two-thirds of their token budget on unreferenced code. Grounded in reinforcement learning and sequential decision processes, we formulate cost-aware agentic scanning across two graphs: casting context retrieval as information acquisition over an exploration graph, and cascade termination as an optimal stopping policy over an adjudication graph. Along the first, anchoring judges strictly to candidate-cited evidence slashes judge inference costs by 44% to 58% across repository-disjoint replications. Along the second, sequential cost-aware stopping halts adjudication on a quarter of production runs, cutting their cost by 25.9% (15.4% across all traffic), while preserving 99.86% of final decisions and 99.2% of verified teacher keeps. Crucially, our evaluations reveal that cost optimization is not a free lunch: cited-file anchoring shifts verdicts beyond the judge’s own repeat variability, while agentic navigators fail to outperform character-capped random retrieval. This study delivers the first end-to-end empirical decomposition of cost-fidelity dynamics in deployed agentic review, proving that context pruning, policy-driven stopping, and security correctness cannot be treated as interchangeable optimizations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.