Yesterday's Facts, Today's Violations: Schema Tightening in LLM-Written Knowledge Graphs
Abstract
LLM agents increasingly write into the knowledge graphs that later ground their answers, and validating each write against a schema is already production practice. In every such system the schema is monotonic: it is frozen or only extended, so data that once passed validation keeps passing. The constraints users ask for run the other way. Writing for the set of graphs that validate against shapes , “every contract must carry at least one document” is a tightening, a change with . It retroactively invalidates data that has not changed, along with answers already delivered from that data. Incremental SHACL validation does not cover this case, because it assumes that the data changes while the shapes stay fixed. We treat a constraint change as a first-class, versioned event. After a policy gate and an impact preview, the system revalidates the whole graph and emits each newly broken fact as a conflict event that records the shape version that judged it and the agent behavior that wrote the value; the value is then repaired, or the event is deferred with the violated shape as its reason. No existing benchmark specifies what breaks under a tightening or what the correct repair is, so we build one by inversion: we delete known values, inject tightenings, and diff the validation reports, and each deleted value becomes the gold repair. On QUDT and LUBM, a single tightening takes the report from to violations without touching the data. Our full loop clears of the violations it attempts but repairs only correctly, and the gap between the two widens from to pp as repair prompts get stronger. A system that reports validity alone therefore overstates its repair rate and silently corrupts the graph it is meant to protect.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.