PolitiBench: A Structured Evidence Benchmark for Post-Training LLM in Fact-Checking
Abstract
Automated fact-checking with large language models (LLMs) is dominated by frontier closed-source models, and prior attempts to specialize smaller models have not closed the gap yet. We ask whether the articles written by professional fact-checkers can serve as post-training data that teaches smaller LLMs to fact- check. We collect, process and convert 26,101 real-world PolitiFact fact-checking data into training instances that pair a contextualized claim with the fact-checker’s analysis, stripped of verdict-revealing sentences, and with evidence extracted from fact-checking articles and organized into supporting facts, weakening facts, and missing context. Our post-training of Qwen3-4B and Qwen3.5-9B, with Group Relative Policy Optimization (GRPO) and our proposed reward function, raises six-class verdict accuracy over the base checkpoints by 5.6% and 7.4%, respec- tively, with particularly large accuracy gains on difficult ambiguous classes (e.g., Half True from 6.51% to 44.79% and Mostly True from 11.02% to 29.02%). In the setting that mirrors retrieval-augmented fact-checking in practical deployment, where only the claim and formatted evidence are available, the post-trained 9B model exceeds the best among the 5 frontier models, GPT-5.6 Luna, by 1.3%. When sufficient evidence is available, our post-trained models improve the accu- racy of their base models by 7.1% and 5.8%, respectively, once again reaching the same level as GPT-5.6 Luna. The gains transfer to 11 external fact-checking eval- uation cases, where our post-trained models outperform their base checkpoints in 80% of cases, and match or surpass frontier models on 8 of the 11 benchmarks under zero-shot settings. This verifies that the post-training on our datasets boosts the general fact-checking ability of small models. We release PolitiBench, the processed corpus, as a resource for both evaluating and specializing LLMs for evidence-grounded fact-checking.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.