acceptodds
Under review as a conference paper at ICLR 2027

Hallucination Mitigation via Directional self-guided, Sample-Selective Contrastive Learning

Abstract

Hallucination in large language models remains a major obstacle to reliable deployment. Existing fine-tuning and preference optimization methods broadly update the model, potentially perturbing otherwise reliable continuation preferences, while treating preference pairs largely uniformly. We instead formulate hallucination mitigation as self-guided selective contrastive post-training, suppressing a model-generated incorrect continuation only when it remains competitive with the correct one. The method requires no expert annotation or manual data curation and adaptively activates updates only on hallucination-active examples, with the selected fraction varying across training distributions. Despite this selective update scheme, it consistently outperforms/it is comparable to full-data SFT and standard preference-optimization baselines under cross-dataset evaluation, yielding a stronger repair–damage trade-off. These results suggest that targeting hallucination-active examples can be effective, while enabling a lightweight and scalable self-guided framework that automatically generates, filters, and corrects its own hallucination cases.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.