A Graph-theoretic Approach to Privacy and Robustness in LLMs
Abstract
Large language models (LLMs) face security risks from both unintended disclosure and adversarial manipulation: they may reproduce sensitive training content, leak membership information, and exhibit substantial performance degradation under small targeted prompt perturbations. Although these security threats differ in their manifestation, they are influenced by how an LLM internally routes attention across tokens. Existing approaches lack a common, training-free framework to study LLM security by analyzing and intervening on this internal process. We adopt a common graph-theoretic lens, representing the causally masked attention map of a decoder-only LLM as a directed acyclic graph, with tokens as nodes and attention scores as weighted edges. Using this representation, we introduce two complementary inference-time approaches for analyzing LLM security vulnerabilities: an edge-centric intervention on attention weights for privacy and a node-centric intervention for context robustness. First, AttentionSwap identifies structural agreements in dominant edges of a primary model and an auxiliary model before swapping their weights to disrupt verbatim memorization while preserving utility on the downstream tasks. Second, BotDrop identifies topologically critical tokens and prunes them to stress-test prompt sensitivity. Extensive empirical evaluations demonstrate the effectiveness of this shared graph-based perspective for addressing different LLM security threats through our novel training-free, inference-time interventions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.