Causal Contracts: Constraint-Verified Decoding for Hallucination-Resistant LLMs
Abstract
Large language models generate fluent and confident text, but their training does not explicitly enforce causal validity. This limitation is critical in high-stakes domains, where residual hallucinations often involve incorrect causal or temporal relations rather than missing knowledge. We therefore define hallucination relative to a domain structural causal model (SCM). A claim is considered inconsistent when its associational, interventional, or counterfactual relation is not supported by the corresponding causal structure. We formulate generation as a Lagrangian-relaxed constrained optimization problem and introduce a projected primal-dual procedure. Under standard smoothness assumptions, we establish a non-asymptotic stationarity guarantee. We also derive a PAC-style bound on residual constraint violations as a function of verifier queries. Finally, we define an evaluation protocol spanning clinical QA, scientific explanation, and causal reasoning. The protocol combines causal-consistency verification, human annotation, inter-annotator agreement, and paired significance testing. Preliminary experiments on real LLM outputs demonstrate the feasibility of graph-based causal verification and identify causal claims that conventional factuality measures can miss.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.