Domain-Grounded Counterfactual Augmentation for Mitigating Knowledge Conflicts in LLM Fine-Tuning
Abstract
Domain-specific fine-tuning of pre-trained LLMs exposes a persistent failure mode: existing parametric knowledge persists as a dominant inductive bias, so models produce confident, internally consistent predictions that remain inconsistent with the target domain even after supervised fine-tuning on it. Such failures are not fully captured by label noise or distributional shift in the conventional sense; they reflect a domain-conditioned knowledge-priority problem, in which generalizations valid under the original training distribution are overridden by more specific target-domain facts, which creates a need for targeted supervision and consistency across domain-preserving views. We introduce CIDA (Conflict-Immunized Data Augmentation), a framework that applies counterfactual invariance to this misalignment. CIDA identifies samples prone to conflict via distributional consistency under stochastic decoding, constructs knowledge-conditioned rewrites intended to vary prior-triggering surface cues while preserving the target answer, and optimizes a composite objective coupling conflict-aware sample weighting with an invariance regularizer on predictions over original and counterfactual pairs. Evaluated on an internal domain-specific dataset and public benchmarks covering diverse domains, CIDA outperforms competitive baselines from data filtering and sample reweighting paradigms. Gains concentrate on high-conflict samples, while low-conflict performance remains comparable to standard fine-tuning. Sample- and update-matched comparisons separate additional training exposure from the augmentation pipeline, while a same-data ablation identifies the additional benefit of paired invariance. Together, these comparisons support retrieval-grounded augmentation and prediction consistency as complementary contributors to the gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.