acceptodds
Under review as a conference paper at ICLR 2027

Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning

Abstract

Reasoning models can keep reasoning even when they can already give the correct answer. We call these additional, post-commitment steps reasoning theater. We introduce ProFIL (Probe-Filtered Reinforcement Learning), a drop-in extension to Group Relative Policy Optimization (GRPO). Prefix counterfactuals identify post-commitment steps; a lightweight probe is trained once on activations of a base model and then frozen. During RL, the current policy produces rollouts, while a separate frozen copy of the base model teacher-forces the same rollout text to supply the probe activations. High-theater rollouts receive zero reward and zero policy advantage. Across GSM8K, LiveCodeBench, ToolUse, and MMLU-Redux and two model architectures, ProFIL reduces post-commitment theater by 11-100%, raises faithful fraction (including +24 percentage points on LiveCodeBench under an independent GPT-4.1 judge), and shortens chains by 4-19% in three of four domains, while preserving or improving the reported task-accuracy metric. A matched length-penalty baseline worsens theater, isolating commitment detection from generic compression. Frozen-probe audits, an independent pairwise judge, and inference-time steering controls further support the result and address the RL-obfuscation concern. Probe weights, training configurations, evaluation code, and rollout caches are released for all four domains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.