GIF: Locally Sound Geometric Information Flow Control for LLMs
Abstract
Large language models increasingly mediate interactions between sensitive data, untrusted inputs, and privileged actions, creating security and privacy risks in agentic systems. Information Flow Control (IFC) offers a promising defense, but existing approaches lack a principled way to reason about information flow through the model itself and therefore suffer from severe taint explosion. We introduce Geometric Information Flow (GIF), a semantic framework that uses the LLM's Jacobian and local output geometry to measure how input spans influence model outputs. GIF upper-bounds the mutual information induced by local input perturbations and admits scalable estimation via automatic differentiation and low-rank approximation. Across integrity and confidentiality benchmarks spanning prompt injection and privacy leakage, GIF consistently outperforms token-level attribution baselines and, when paired with lightweight declassifiers, matches or exceeds direct LLM-as-judge baselines while using up to lower model inference cost. GIF also transfers across model scales and families, including from surrogate models up to smaller than their targets. Our results establish geometric information flow as a principled and practical foundation for scalable IFC in LLM systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.