acceptodds
Under review as a conference paper at ICLR 2027

JAAL: A Benchmark of Hinglish Scam Conversations

Abstract

Scam calls are inherently conversational, yet publicly available fraud datasets provide limited coverage of multilingual and code-switched conversational interactions, particularly in Hinglish. We introduce JAAL, a synthetic Hinglish scam-call corpus and benchmark containing 9,677 multi-turn conversations across eight fine-grained scam categories. The conversations are generated through interactions between caller and receiver agents powered by Gemma 3 27B. We systematically characterize the corpus in terms of class distribution, conversation length, language composition, vocabulary, subword fragmentation, and duplicate and near-duplicate structure. We further establish a benchmark spanning classical classifiers, neural models, pretrained language encoders, and zero-shot large language models. On a stratified held-out synthetic test set, supervised models achieve strong performance, with RoBERTa reaching 92.33% Macro-F1. To assess generalization beyond the synthetic benchmark, we additionally evaluate the same models on 77 independently sourced YouTube scam conversations. The external evaluation reveals substantial differences from the synthetic data in language composition, lexical novelty, conversation length, and subword fragmentation, providing a complementary stress test for models trained or evaluated on the synthetic corpus. JAAL is designed to support reproducible research on conversational scam detection in Hinglish and to facilitate systematic study of model performance under differences between generated and independently sourced conversational data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.