acceptodds
Under review as a conference paper at ICLR 2027

HAFT: Adversarial Robustness Benchmarking for Hindi Fact-Checking with Adaptive Attack Selection

Abstract

Hindi is the world's third most spoken language (over 610 million speakers), yet automated fact-checking systems in this ecosystem face acute risks from unverified digital content and forwarded misinformation without provenance. We introduce HAFT, a dedicated Hindi adversarial robustness benchmark comprising 1,120 claims across six news domains and 22 controlled attack mechanisms, yielding 24,640 quality-gated evaluations. Under uniform automated verification and linguistic quality control, we uncover substantial asymmetry across attack families: under authentic evidence retention, a single injected passage flips 57.6–58.5% of correct verdicts (reproducing in 199–200 of 200 audited cases under an independent open-weight 70B model), whereas character perturbations remain below 3.7%. We then evaluate budgeted defensive auditing under a constrained testing budget of 5 of 22 attack attempts. A static ranking of evidence attacks discovers 86.71% ± 1.16% of vulnerabilities, while an adaptive claim-conditioned policy achieves 82.16% ± 2.02% (vs. 40.84% for random attempts and 72.69% for an online bandit), trading a small discovery margin to expand diagnostic exploration across 16 distinct mechanisms. Finally, on a 1,142-node heterogeneous claim–attack graph, GraphSAGE reaches 0.865 AUROC under random splitting but degrades under strict inductive Leave-Attack-Out (0.485 AUROC), and relational graph embeddings yield statistical parity in policy performance (80.72% vs. 82.16%, p = 0.443). HAFT provides reproducible empirical foundations for adversarial robustness and budgeted auditing in low-resource automated fact-checking.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.