SecFlow: Learning to Interpret and Use Security Evidence in Language Models
Abstract
Understanding the surface meaning of security evidence does not ensure that models interpret its security implications or use it correctly. We investigate how reliably large language models (LLMs) understand security facts and use evidence to guide decisions. Using vulnerability information extraction and patch localization as complementary study settings, we identify limitations in both capabilities across seven commercial LLM systems. Our audit attributes 73.28% of incorrect localization runs to dominant reliance on surface cues and documents cases where systems inspect the correct patch but fail to select it. We introduce SecFlow, which represents textual and program evidence as source-linked behavior flows encoding operations, conditions, and effects. It evaluates these flows against the vulnerability under analysis and injects the routed memory into later language-model layers through gated attention to guide answer generation. Built on Qwen2.5-Coder-7B, SecFlow improves both security-fact understanding and evidence-guided decision-making. Across the two study settings, it increases grounding F1 by 16.51 percentage points and achieves 65.24% Top-1 accuracy and 95.92% Hit@5 in candidate selection, surpassing the strongest evaluated commercial baseline by 40.13 percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.