acceptodds
Under review as a conference paper at ICLR 2027

Auditing Cross-Lingual Chain-of-Thought Probes with Leakage Controls

Abstract

Chain-of-thought (CoT) monitors read a model's written reasoning to catch misbehavior. Outside English, that text often fails: the model can follow a planted wrong hint and never mention it. A common fix is to look inside the model instead. Train a linear probe on English hidden states, then reuse it in other languages, assuming the internal "about to follow the hint" state is shared. We test that idea on a misleading-hint task in eight languages and four scripts, using three open models: Qwen3-8B (thinking), plus Apertus-8B and EuroLLM-9B (instruct). We also extend the prediction check to a same-content Azerbaijani Latin/Perso-Arabic script swap and to a wider set of thinking models across size and family (Qwen3 1.7B/32B, R1-Distill-Llama-8B, Phi-4 Reasoning). After holding out question IDs, the English probe still predicts hint-following at post-CoT (mean non-English AUROC 0.69–0.79), and removing a linear language signal barely changes transfer. But much of what it reads is already written: filtering early answer letters leaves Qwen3-8B near 0.61, while some instruct-model scores fall to chance; cutting the end of the reasoning drops all three late-probe means to 0.52–0.56. And the same English directions do not control the behavior. Adding or subtracting them during generation is no better than a same-size random direction in English, Turkish, Persian, or Swahili; stronger and multi-layer injections also fail to give reliable signed effects. So an English probe can monitor hint-following across languages, but the directions we test cannot steer it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.