When Models Struggle with Negation: Learning from Sound and Unsound Reasoning
Abstract
Direct Preference Optimization (DPO) is only as effective as the preference data it learns from. As human annotation is expensive, synthetic preference pipelines increasingly rely on large language models (LLMs) to generate candidate responses and on LLM judges to identify preferred ones. Yet this pipeline implicitly assumes that the generator can already produce sufficiently reliable responses for the target capability. When models struggle with that capability, preference construction becomes circular: it presupposes the very capability that post-training is intended to improve. We study this problem through negation, a pervasive linguistic phenomenon that remains challenging for LLMs across multiple tasks. We introduce PACE, a framework that constructs preference data without requiring models to first solve the target problem reliably. Given known correct and incorrect answers, PACE elicits corresponding justifications that reflect contrasting sound and plausible but unsound reasoning. PACE further uses pairwise comparisons with Bradley-Terry aggregation to rank candidate justifications within each answer-conditioned set while mitigating positional bias. Across five negation-focused benchmarks spanning question answering, natural language inference, commonsense reasoning, and information retrieval, DPO training on approximately 5K PACE preference pairs yields gains across four Llama and Qwen backbone LLMs, with statistically significant improvements in 17 out of 20 model-benchmark settings. Controlled experiments show that removing justifications conditioned on incorrect answers consistently weakens performance, demonstrating that plausible but incorrect reasoning provides an important learning signal. Our results suggest that reliable preference construction is a critical prerequisite for effective preference optimization when models struggle with the target capability. To our knowledge, we release the first preference dataset for negation understanding, along with our implementation at https://anonymous.4open.science/r/NEG.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.