Reasoning over Broken Joins: How LLMs Handle Integrity Violations in Multi-table Inputs
Abstract
Large language models are increasingly used to query multi-table relational data. The correctness of such queries depends on integrity constraints such as key uniqueness and referential integrity, yet these constraints are frequently violated in practice. On such corrupted inputs, a model should detect the violation and disclose it rather than answer silently. Existing evaluations on such corrupted inputs score only the final output. They therefore cannot distinguish a model that misses a violation from one that detects it but answers without disclosing it. We introduce RELIC, a dataset of approximately 16,000 multi-table instances in which six violation types are injected at controlled severities into integrity-verified synthetic table pairs. We define the silent failure rate (SFR) as the fraction of corrupted instances that a model answers without disclosing the violation. We further decompose silent failures by whether the model detects the violation in its reasoning trace. SFR exceeds 84% for all 20 models and 97% for all three frontier models despite clean accuracy of up to 98.3%. Detection increases with model capability and reasoning budget while disclosure does not. Prompt-level interventions elicit disclosure of duplicated primary keys from reasoning models but leave most orphaned foreign keys undisclosed. Model scaling and prompt-level interventions therefore fail to mitigate silent failures consistently across violation types. RELIC enables fine-grained attribution of silent failures to detection or disclosure for each violation type, facilitating the development of LLMs that handle integrity violations reliably.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.