acceptodds
Under review as a conference paper at ICLR 2027

Answer Set Consistency of LLMs for Question Answering

Abstract

Large Language Models (LLMs) can contradict themselves when answering factual questions, especially when asked to enumerate all entities that satisfy a question. We formalize such self-contradiction as answer-set inconsistency: given two enumeration questions whose answers satisfy a set-theoretic relation (equivalence, disjointness, containment, etc.), the LLM generates responses violating the relation. To diagnose this phenomenon, we create Answer-Set Consistency Benchmark (ASCB), a benchmark dataset comprising 600 tuples of enumeration questions (2400 questions in total) over which a variety of set-theoretic relations hold, and propose related metrics to quantify answer-set inconsistency. Our evaluation of 18 state-of-the-art LLMs reveals pervasive inconsistency across models, even when the LLM can identify the correct relation. We also analyze likely sources of these failures and use relation-aware prompting as a lightweight exploratory intervention. Making the relation explicit often improves consistency, but gains are uneven and can be offset by abstention or prompt misinterpretation. Our results position answer-set consistency as a complementary dimension for evaluating LLM-based question answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.