NetTrace: Evidence-Bounded LLM Reasoning for Network Traffic Investigation
Abstract
Tool-using language models can cite actual observations yet draw conclusions those observations do not support. In network investigation, support depends on the evidence set and a claim’s entity, time, traffic layer, quantifier, and strength. We formulate evidence-bounded traffic investigation and introduce NETTRACE, combining structured traffic inputs, claim–evidence supervision, and a separately trained scope-aware verifier with conflict checks and revision. On NetTraceBench’s 4,800-case test split, NetTrace achieves 98.96 macro-F1 and a 2.84% non-entailing-link rate in the supplementary batch. A same-scale text verifier with matched deterministic scope checks achieves 86.86 and 5.62%, respectively. The F1 difference is 12.10 points; separate controls show advantages under fixed observations and matched answer coverage. Counterfactual pairs, jointevidence tests, stricter splits, and an independent human audit examine the support mechanism under controlled interventions and distribution shifts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.