acceptodds
Under review as a conference paper at ICLR 2027

Reading Is Not Simulating: Calibrating Chain-of-Thought Interventions on an Executable Cognitive Model

Abstract

Chain-of-thought interventions probe answer dependence on written reasoning. Interpreting that dependence as consequence prediction requires a reference for what an edit should change. We build one from an executable cognitive process: an ACT-R-derived driver runs in a closed-loop simulator, and 1.5B–7B language models are fine-tuned to generate its templated state traces before predicting outcomes. Shared-randomness re-executions supply mechanism-specific counterfactuals; fresh seeds at the same specification supply empirical consequence references. At 1.5B, whole-trace replacement reaches 74% answer-direction agreement on outcome-moving pairs, yet donor replacements without the corresponding mechanism intervention yield comparable or higher agreement against their own outcome changes. Edited states persist in 99–100% of evaluated continuation clauses where reference traces disagree. In the matched-boundary assay, all three students fall below the specification-matched majority for lane-consequence accuracy on the hands mechanism; the 7B student also falls below arm-wise modal prediction and re-execution agreement. Separately, hands answer-direction agreement is 79–85% with consequence-supplied or full-trace inputs, versus 7–52% after state-only continuation. Reference decision rules change which shortfalls are supported, and a layout audit exposes termination leakage. In this setting, we empirically distinguish following supplied traces from predicting consequences of edited states, and provide an executable protocol that makes this distinction auditable through explicit information boundaries, decision rules, and evaluation populations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.