SC-CoIT: Self-Corrective Interleaved Chain of Thought via Bidirectional Cross-Modal Recovery
Abstract
Unified multimodal models (UMMs) can generate Interleaved Chain of Thought (Interleaved CoT) traces that alternate between language and visual thoughts. Yet correct-only supervised fine-tuning provides no recovery target when an intermediate node fails. We identify two cascade failures: a Draw Error, where the text expert proposes a valid action but the image expert misexecutes it, and a Policy Error, where the text expert proposes an action inconsistent with the task or environment and the image expert follows it without verification. From ground-truth Interleaved CoT traces, we construct a bidirectional error-recovery training dataset and propose (SC-CoIT), which trains the modality experts to check each other's outputs and recover from cross-modal errors. For Draw Errors, the text expert grounds the discrepancy and directs a same-step redraw; for Policy Errors, the image expert emits a visual reflection signal, prompting the text expert to regenerate the step. To adapt both recovery branches to failure modes that evolve during training, iterative error replay samples new errors from the current checkpoint on validated prefixes and adds them to subsequent training. Together, these mechanisms improve the model's ability to inspect generated visual states, reject invalid actions before execution, and recover from both error types within the same Interleaved CoT trajectory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.