acceptodds
Under review as a conference paper at ICLR 2027

A Few Contradicted Examples Silently Erode Transfer to Held-Out Formats and Languages

Abstract

Instruction tuning is valuable because it transfers: tuned models follow instructions that specify answer formats, response templates, schemas, or output languages even for targets unseen in training, but we show that this transfer erodes silently when a small fraction of training responses contradict their own instructions. Holding responses fixed and changing only instruction–response consistency, 5% contradiction lowers held-out compliance by 19–62 points across seven models from 0.5B to 8B, with full contradiction driving transfer to zero. In Qwen2.5-1.5B, just 2% contradiction reduces held-out answer-format compliance from 68.3% to 47.7%, template compliance from 51.6% to 27.8%, and language compliance from 90.4% to 81.0%, while trained-instruction compliance changes by at most 1.4 points; similar drops appear in SmolLM2-1.7B and Qwen2.5-7B. Removing contradicted examples restores transfer, whereas removing the same number of random examples does not. The effect also appears in public data: removing the 2.45% of xLAM tool calls that violate their schemas eliminates 43% of held-out schema errors, and repairing natural constraint violations in Tülu 3 improves held-out compliance across every seed without affecting trained-condition metrics. We explain why the failure is difficult to detect: at the fine-tuning optimum, trained-condition quantities depend only on contradiction rate ρ, with validation loss increasing by −ln(1−ρ) nats (0.02 at 2%), while held-out behavior remains unconstrained. Standard data-quality signals—including loss, IFD, Superfiltering, and an LLM judge—identify contradictions poorly (AUROC ≤ 0.82), whereas instruction-specific verifiers detect all 2% contradictions with no false positives and a label-free likelihood-swap test achieves AUROC 0.997–0.999; removing its top 5% restores mathematics transfer to the clean level (67.8% versus 66.9%).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.