acceptodds
Under review as a conference paper at ICLR 2027

RASA: Role-specific Acoustic–Semantic Alignment for Telecom Fraud Detection

Abstract

Distinguishing fraudulent calls from legitimate conversations requires interpreting callers' claims and requests within the surrounding conversational context. Acoustic behaviors such as disfluency, speech rate acceleration, and frequent interruptions can supplies crucial acoustic evidence when assessing deceptive intent. However, transcript-based large language models may overlook these cues, while direct audio input into large audio-language models (LALMs) does not guarantee their adequate preservation by recognition-oriented encoders. So that we propose Role-specific Acoustic–Semantic Alignment (RASA), which combines a recognition-oriented speech encoder with a reconstruction-oriented neural audio codec. Separate Fixed-Length Query Adapters aggregate their representations into compact soft-token sequences for joint reasoning within an LLM. Using transcript- and script-informed teacher supervision from ParaScam, our synthesized call dataset, role-specific alignment establishes a semantic anchor and trains the acoustic pathway to contribute acoustic evidence. The aligned model supports downstream adaptation using only audio and task labels, with no transcripts required at inference. On TeleAnti, RASA achieves 98.13% fraud-detection weighted F1 and 88.61% average weighted F1 across fraud detection, fraud-type identification, and scene classification. Under a 500-call downstream adaptation setting, RASA improves fraud-type F1 by 27.37 percentage points over its semantic-only variant, supporting the value of explicitly aligned acoustic evidence for fine-grained fraud recognition.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.