Answer Set Consistency of LLMs for Question Answering
Abstract
Large Language Models (LLMs) can contradict themselves when answering factual questions, especially when asked to enumerate all entities that satisfy a question. We formalize such self-contradiction as answer-set inconsistency: given two enumeration questions whose answers satisfy a set-theoretic relation (equivalence, disjointness, containment, etc.), the LLM generates responses violating the relation. To diagnose this phenomenon, we create Answer-Set Consistency Benchmark (ASCB), a benchmark dataset comprising 600 tuples of enumeration questions (2400 questions in total) over which a variety of set-theoretic relations hold, and propose related metrics to quantify answer-set inconsistency. Our evaluation of 18 state-of-the-art LLMs reveals pervasive inconsistency across models, even when the LLM can identify the correct relation. We also analyze likely sources of these failures and use relation-aware prompting as a lightweight exploratory intervention. Making the relation explicit often improves consistency, but gains are uneven and can be offset by abstention or prompt misinterpretation. Our results position answer-set consistency as a complementary dimension for evaluating LLM-based question answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.