Logic-Aware Policy Hallucination Detection through Text-to-Graph Encoding
Abstract
Translating policy text into an executable program can omit conditions needed for the correct decision. As a result, policies that should produce different decisions for the same question and context may produce identical outcomes after translation. This loss of semantic sensitivity contributes to hallucination when the system asserts an unsupported conclusion. Successful symbolic execution alone cannot establish source support, so automatic detection must connect source meaning with the completed claim. We learn continuous, query-conditioned graph embeddings to recognize such failures in supplied verifier runs. A graph-only branch predicts candidate execution, while a source-aware branch aligns policy units with program nodes and scores each asserted claim without requiring a comparison partner. A paired branch uses symmetric and directional features to classify changes in semantic sensitivity. Logic-derived supervision trains these differentiable representations without sending gradients through the compiler or symbolic evaluator. Source-grounded repair and replay assess contribution separately. On 10,845 held-out controlled pairs from 600 policy families, sensitivity accuracy is 63.89%. Individual-claim detection achieves 81.24% F1 and 91.88% average precision on 4,635 evaluable assertions, against a 65.46% constant-score baseline. On a structurally broader benchmark, average precision was 56.56% among scored claims; a subsequent capacity check accommodated all 3,974 evaluable claims and yielded 56.44%. These single-seed results establish a detection signal in controlled logical settings alongside substantial transfer failures. Outcome recoverability, executable fidelity, and source-supported claims remain distinct.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.