OWNERSHIP IS THE MISSING DECLARATION: DECIDING RELATIONAL TARGET LEAKAGE FROM THE SCHEMA
Abstract
Target leakage is treated as a statistical property of data: a column is leaky if a probe trained on it predicts the target too well, so no column can be judged before it is read. We ask instead which features the declared schema forces to determine the target on every database it admits, and what any schema-only method must leave undecided, where a task is a tree of many-to-one lookups over keys, foreign keys and copies. On a fragment allowing composite, cyclic and self-referencing foreign keys, a whole-key closure decides this soundly and completely, in time near-linear in the copy-closed view, and every negative verdict carries a counterexample database. One further SQL-declarable construct makes the question undecidable for feature sets, and targets aggregated over future rows, such as churn, are outside the fragment. Four such mechanisms cost models up to 22.4 accuracy points and 54.5 macro-F1, and the closure abstains on every one. Across 5,094 WikiDBs databases, 119 CTU databases and 23 benchmark tasks, keys and foreign keys alone certify no non-identifier leak at all; across 22,989 public schemas they certify 124 of 1,137,685 targets, every one a key cycle. Two open-source ERPs do declare which table owns a value, yet neither dump carries it, and one never emits a database-level foreign key at all. Blind language-model declarers recover 0 of the 90 columns designers removed by hand: the method certifies what a declaration forces, not leaks found unaided. Checking a declaration on a sample is cheap and carries a bound. On one RelBench task a name-blind generator proposes the declaration and ten sampled rows accept it; the closure certifies the removal, and re-inserting the column raises accuracy by 28.0 points. Given an ownership declaration, the rest of the fragment is decided before a row is read.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.