acceptodds
Under review as a conference paper at ICLR 2027

Minimax Learning from Corrupted Chain-of-Thought Supervision

Abstract

Chain-of-thought (CoT) supervision has been used to train reasoning models, but the intermediate reasoning steps and final answers provided for training can be corrupted. We study learning from corrupted outputs of a deterministic autoregressive teacher model that belongs to a finite candidate class. We establish matching upper and lower bounds on the number of prompt-level training examples needed to achieve a target clean prediction error, in the worst case over classes of a given size. For final-answer supervision, we consider Huber contamination: with probability , each answer is replaced by a draw from an arbitrary unknown distribution. We show that empirical risk minimization (ERM) is minimax optimal and characterize how the contamination probability increases the required sample size. For full-CoT supervision, tokens are corrupted independently after the teacher generates the complete sequence, and prediction error is measured by the expected fraction of incorrect tokens. When the token corruption probabilities are known and distinct clean tokens induce distinct observation distributions, maximum likelihood is minimax optimal. Its sample complexity is governed by the distinguishability of the corrupted observations, accumulated over the token positions on which candidate rollouts disagree. For unknown Huber token contamination with probability , minimizing token disagreements between candidate-generated sequences and observed sequences remains minimax optimal, without knowing either the corruption probabilities or . These results quantify the statistical cost of corrupted supervision and show that channel knowledge can sharpen guarantees for particular corruption mechanisms, although it is unnecessary for optimal worst-case performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.