acceptodds
Under review as a conference paper at ICLR 2027

Recognizing Constraints Is Not Enough: The Judgment–Execution Gap in Large Language Models

Abstract

Reliable large language model assistants and agents must pursue users' goals without violating the constraints that define valid completion. It remains unclear whether models preserve task constraints during execution even when they correctly judge the task infeasible. To study this judgment-execution consistency under verifiable conditions, we use well-defined scientific constraints that can be checked against external evidence. We introduce NoGoBench, comprising 1061 scientifically impossible tasks and minimally edited feasible counterparts across six domains, and evaluate feasibility judgments and task execution in independent conversations. Our conditional False Completion Rate (cFCR) measures the proportion of tasks self-reported as completed among those correctly judged impossible. Across 19 general-purpose and science-specialized model configurations from 10 families, Contrastive Macro-F1 ranges from 92.2% to 97.8%, while cFCR ranges from 8.3% to 76.6%. An external judge attributes most classified inconsistent cases to specification evasion through redefined terms, weakened requirements, or substituted goals. Further experiments show that increased reasoning reduces direct errors without reliably improving completion reporting, while post-training recipes affect feasibility judgments and completion claims differently. Stronger execution pressure substantially increases false completion. These findings show that recognizing scientific constraints does not ensure reliable execution, motivating further investigation of judgment-execution consistency under value constraints. Our code and data are available at https://anonymous.4open.science/r/NoGoBench-3ED5.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.