ROWGATE: DETERMINISTIC ROW-LEVEL ACCESS CONTROL FOR LLM QUERY AGENTS OVER FEDERATED SOURCES
Abstract
Large language model agents that turn natural language into SQL are now being put in front of enterprise data, where the access policy has to hold at the row level and the data usually cannot leave the systems it already sits in. The common practice is to write the policy into the prompt and trust the model to follow it. We measured how well that works, and it breaks in two directions at once. On a 150-probe adversarial benchmark covering five RBAC roles and ten models from four vendors, refusal rates on identical inputs run from 0.0% to 72.0% and schema fidelity from 54.4% to 100%. The two do not move together. The most cautious model turns away 72% of perfectly ordinary business queries; another turns away none and matches it on fidelity; a third turns away none but reaches into unauthorized tables 45.6% of the time. Even when the caller’s scope is supplied as context, models leave the mandatory row filter off 23% to 32% of the SQL they do produce, and rewording the prompt moves that number. RowGate takes the model out of the trust path instead. Policy is enforced at two deterministic layers built over a four-graph metadata substrate: unauthorized tables and columns are stripped from the prompt before generation, and the caller’s mandatory predicates are injected into every scope of the parsed SQL before execution, including the per-source sub-queries of a federated plan. We state the guarantee as a partition over outcomes and prove it under three conditions. Prior work checks generated SQL statically; we instead run the rewritten queries against live SQL Server and PostgreSQL instances holding the real AdventureWorks database, and score exact result-set equality against an authorized-only ground truth. Over more than 700 executed queries, covering complex query shapes, a 1,000-table catalog, irregular naming, and a cross-engine join that falls from 121,317 rows to the 5,976 authorized ones, RowGate returns 0 unauthorized rows every time, and exactly the authorized rows in all but 9 long-tail cases, which fail closed rather than leak. Executing the queries also turned up four row-level leaks in our own rewriter that a static predicate check would have passed. We then overlay policies mechanically onto the Spider benchmark and run 495 queries over 17 databases the rewriter has never seen: all exact, 0 leaks, against 192 that a rewriter injecting nothing would leak on. Enforcement matches database-native row-level security on a single engine, and extends where that cannot reach, across engines and ahead of generation. The rewrite is a pure function of the query and the policy; its output is byte-identical at 33 tables and at 1,000. We release the benchmark, the invariant taxonomy, and the harness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.