acceptodds
Under review as a conference paper at ICLR 2027

Readable Constraints, Unreliable Decisions: A Constraint-Instantiation Lens on Language Model Errors

Abstract

Natural answers expose a model’s final judgment, but may not express all task-relevant information recoverable from its internal states. We study whether suchrecoverable information tracks the truth required by the current input and whetherdisagreement between recoverable truth and the model’s decision can diagnoseerrors. In a verifiable Fact–Rule–Action (FRA) setting, we distinguish task-definedlocal truth, hidden-state readouts, and natural responses. Correct local truth remainsrecoverable under answer errors, and frozen readouts systematically track truth-changing edits to facts, rules, and action arguments across new worlds. On aheld-out controlled cohort, readout–answer disagreement strongly ranks natural-answer errors. We further extend the same truth-readout principle to natural QA,where the model’s realized answer instantiates the proposition to be evaluated. Theresulting error-oriented post-answer readout, FRA-ER, improves error ranking overother baselines on several tasks across several model families.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.