acceptodds
Under review as a conference paper at ICLR 2027

Cusp: Evaluating reasoning on the boundary of solvability

Abstract

Language models have demonstrated notable progress on a variety of reasoning and mathematics tasks. We consider reasoning in the context of set theory and propose a benchmark for evaluating the ability of language models to reason about sets under constraints. We programmatically construct and verify our problem instances. Datasets in our benchmark comprise problem instances with minimally sufficient constraint sets that places them on the boundary of (unique) solvability. We evaluate the ability of language models to perform two separate tasks: (1) generate an assignment of elements to the sets in the problem such that all constraints are satisfied and (2) state how many solutions exist to the problem. We find that on the boundary of solvability, accuracy on these reasoning tasks can vary substantially depending on two levers that we pull simultaneously: the complexity of the problem instance and the pragmatics of the prompt. Some of the behaviors we observe are consistent with reports of short-cutting in earlier literature, while others are novel and warrant further investigation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.