acceptodds
Under review as a conference paper at ICLR 2027

MEASURING REASONING REQUIRES LETTING MODELS REASON

Abstract

Logical reasoning in language models is evaluated one inference rule at a time, and often under a prompt that asks for a bare True/False label. We built a matched-arm benchmark to measure what a weak rule costs inside a chain—three arms sharing a four-step derivation and differing only in the form of one rule—and under that answer-only format the result is dramatic: across five frontier models a rule with a negated premise is applied at 0.999 and its contrapositive at 0.209, with GPT-4.1 correct 0 times in 600 attempts while every other step of its chain is at 1.000. Asked for the same contrapositive on its own, GPT-4.1 is correct on 100 of 100 instances and rejects 100 of 100 foils, so the rule is present and is not a surface cue. The natural reading is that chain context destroys a rule the model has. It is the wrong reading. Lifting one sentence from the system prompt, so the model may write its derivation before answering, moves GPT-4.1 from 0.000 to 0.613 on the chain and flattens a depth sweep that had fallen 1.000 → 0.000 over eight derived steps to 1.000 → 0.917 on the same instances. What the answer-only format measures is whether a model can perform an inference without externalising it, which is not the question such benchmarks are taken to ask, and the gap between the two questions can be the entire result. One model is not rescued: GPT-5.2 answers the contrapositive correctly 0 times under every condition we tried—answer-only, internal reasoning tokens, and explicit step-by-step—while applying the negated-premise rule at 1.000 throughout. Code, data and per-run logs: https://anonymous.4open.science/r/lemo-D708/README.md.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.