Evidence Before Assertion: Calibrated LLM-Assisted Verification of Code Changes
Abstract
Large language models can explain code changes fluently, but fluency is a poor substitute for evidence when a reviewer asks whether a route is affected, whether a test actually proves the changed behavior, or whether an unexecuted check is safe to ignore. We present an LLM-assisted verifier for pull requests built around an explicit epistemic contract: a verification claim may be reported as established only when it is grounded in inspectable structural or execution evidence; LLM outputs may propose missing facts, but cannot silently promote them to proof. The system represents behavioral claims as verification obligations, attaches evidence at the strongest scope that can be demonstrated, and degrades attribution rather than confidence when exact ownership cannot be established. We study this design through controlled pull requests on eight real open-source repositories spanning Python, TypeScript, Java, Go, and Rust. The study surfaced seven reproducible failure classes at the boundaries between structural analysis, LLM recovery, test mapping, execution, and reporting. Two cases are particularly informative. First, an apparently intermittent call-resolution failure was shown, after isolating the deterministic subsystem, to be a fully reproducible structural defect masked by occasional LLM recovery; after three generic fixes, the same harness resolved twice and the full pipeline resolved the path in runs. Second, whole-suite execution produced 13 passing and 4 unrelated environment failures for a narrowly scoped change; after wiring mapped-test execution and improving assertion recognition, the system executed exactly one relevant test and attributed a pass to the changed behavior without running the unrelated suite. These results are not a benchmark-scale accuracy claim. They are an empirical systems study of a recurring problem in hybrid LLM tools: correctness failures concentrate at evidence boundaries, and reliable behavior requires explicit provenance, scope-aware evidence ownership, and a calibrated handoff when the environment cannot complete verification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.