Cusp: Evaluating reasoning on the boundary of solvability
Abstract
Language models have demonstrated notable progress on a variety of reasoning and mathematics tasks. We consider reasoning in the context of set theory and propose a benchmark for evaluating the ability of language models to reason about sets under constraints. We programmatically construct and verify our problem instances. Datasets in our benchmark comprise problem instances with minimally sufficient constraint sets that places them on the boundary of (unique) solvability. We evaluate the ability of language models to perform two separate tasks: (1) generate an assignment of elements to the sets in the problem such that all constraints are satisfied and (2) state how many solutions exist to the problem. We find that on the boundary of solvability, accuracy on these reasoning tasks can vary substantially depending on two levers that we pull simultaneously: the complexity of the problem instance and the pragmatics of the prompt. Some of the behaviors we observe are consistent with reports of short-cutting in earlier literature, while others are novel and warrant further investigation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.