acceptodds
Under review as a conference paper at ICLR 2027

One Edit, Three Worlds: Evaluating LLM Reasoning Beyond Answer Accuracy

Abstract

A reliable reasoner must be invariant to changes that do not matter and sensitive to the one change that does. Benchmarks measure the first property well and the second rarely, because every item is scored in isolation. Few prior studies, if any, systematically analyze whether the answer follows the premise that decides it. We propose an evaluation principle that tests both properties together. It builds sibling problems that differ in exactly one premise, and it counts an answer as correct only if the answer changes with that premise and stays the same when the problem is asked again together with its siblings. We call the resulting score verified accuracy. We instantiate the principle as One Edit, Three Worlds (OETW), a benchmark of 1,200 triads, each a base problem with three siblings whose one edited premise leaves the target with no valid answer, exactly one, or two, built on three constraint problems and certified by a solver rather than an LLM judge. The edit is minimal for the model: siblings are byte-identical outside the edited premise, surface-feature classifiers stay near chance, and where probabilities can be read the siblings receive nearly identical ones before reasoning begins. As problems grow, the raw accuracy of seven frontier and open-weight models falls to 33% and their verified accuracy to 2%, where the first is the level of random guessing, and the second means almost every isolated correct answer fails to survive its siblings. What remains is not random guessing but a low-entropy, model-dependent fallback, and neither abstention nor clarification repairs it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.