ROBUST-VERIFY: Adversarial Self-Play with Structured Reasoning for Robust Claim Verification under Corrupted Chain-of-Thought
Abstract
Structured chain-of-thought (CoT) reasoning has become standard in automated claim verification, improving interpretability by requiring a model to articulate intermediate steps before issuing a verdict. However, this approach exposes a previously unaddressed vulnerability that intermediate reasoning steps can be corrupted independently of the underlying claim or evidence, causing verdict errors even when the model possesses sufficient knowledge for a correct judgment. No existing framework provides a mechanism for detecting such corruptions. Zero-shot prompting across four model families (LLaMA-3-8B, Mistral-7B, Qwen2.5-7B, GPT-OSS-20B) achieves Corruption Detection Rate (CDR) near 0.000, confirming that detection requires explicit supervision regardless of model size. We introduce ROBUST-VERIFY, a three-phase training framework augmenting structured verification with an explicit Corruption Check element and training the model to detect and recover from corrupted reasoning through adversarial self-play. Phase 1 performs warm-up training on human-annotated chains with direct inconsistency-detection supervision. Phase 2 trains a Claim Verification (CV)-Polluter and CV-Agent through a policy-gradient method (GRPO) across a six-type corruption taxonomy at three severity levels. Phase 3 introduces Partial Answer Masking (PAM) to concentrate repair supervision on the post-corruption continuation. We present the first explicit evaluation of corruption detection in structured claim verification. Our model achieves consistently strong CDR across HoVer (2/3/4-hop) and FEVEROUS-S, with CDR-L1 ranging from 0.587 to 0.765, a 10x improvement over the best zero-shot baseline while also improving verdict recoverability under several corruption settings. Ablation confirms that PAM is necessary as its removal causes structural collapse (SP = 0.000). Our results show a detection-recovery trade-off between corruption severity levels, motivating CDR-aware reward design as future work.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.