Training Fact-Checkers with Reinforcement Learning
Abstract
Large language models (LLMs) often generate content that contains factual errors, hindering their reliability. To detect these factual errors, we propose FactCheckRL, a new reinforcement learning (RL) based paradigm for training search augmented fact checking agents. With FactCheckRL we train fact checking agents that can flexibly reason and search the web to find factual errors in LLM responses. On two human-verified evaluation sets, FactCheckRL substantially outperforms prompt-based fact-checking pipelines including SAFE, VeriScore, and FaStFact at recovering factual errors identified by human annotators. We further propose to use FactCheckRL as a factuality reward alongside general helpfulness reward models in online RL, showing that this can substantially reduce the amount of factual errors without degrading overall response usefulness. Using this approach, we are able to post-train a model that is more factual than state-of-the-art open models while maintaining comparable helpfulness. We will open-source our code, training data, and models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.