What Role Does Explicit Chain-of-Thought Play in Safety Moderation? A Circuit-Level Study
Abstract
In safety moderation, explicit chain-of-thought (CoT) is often taken as evidence that a model is making careful and trustworthy judgments. Yet it remains unclear whether generating CoT actually helps form the final moderation decision, or whether it mainly verbalizes a decision that is already taking shape internally. To answer this question, we compare direct-answer and full-CoT prompting in a reasoning-based safety moderation model and trace their internal computation at the circuit level. Across multiple benchmarks, we find that the two prompting modes often share a substantial attributed backbone, while disagreement is concentrated in a smaller set of boundary-related features. We also find that much of the model’s latent moderation prediction is already represented at intermediate layers, with later computation increasingly tied to expressing that preference at the output. These results suggest that the final label under full-CoT prompting does not typically emerge entirely from scratch during reasoning generation. Instead, even when direct-answer prompting produces a different final label, its internal state often already reflects support for the label produced under full-CoT prompting. Disagreement between the two modes arises mainly when the internal state under direct-answer prompting remains uncertain, which is also when explicit CoT appears to matter most. Based on this insight, we introduce an uncertainty-guided two-stage moderation pipeline that uses intermediate-layer sparse features for fast prediction and reserves full-CoT reasoning for uncertain cases only. The resulting system reduces reliance on expensive full-CoT inference while maintaining strong moderation performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.