acceptodds
Under review as a conference paper at ICLR 2027

BhashaBench-Agentic: A Multilingual, Multi-Turn Benchmark for Agentic Tool Use

Abstract

An assistant that can look up a government scheme in English proves little about whether it can be trusted to do the same in Hindi. Agentic assistants are already being deployed across Indian public-service and fintech settings, where people write in Devanagari, in romanized Hindi, or in sentences that switch between Hindi and English mid-thought and today's agentic benchmarks were built without that user in mind. We present BhashaBench-Agent to address that gap directly. Rather than another multiple-choice test of factual knowledge, it evaluates whether an agent can act correctly inside a multi-turn conversation: asking when a required detail is missing, confirming before any state-changing action, recovering from a failed tool call, and grounding its answer in what a tool actually returned rather than what sounds plausible. The corpus comprises 4,258 scored episodes built from 82 seed scenarios across 17 tools and 19 languages, spanning agricultural advisory, government scheme eligibility, legal and general-knowledge document work, mandi price lookup, and UPI payments, drawn from four distinct data-curation pipelines : hand-authored, persona-conditioned, mined from BhashaBench-Multi, and mined from de-identified Supreme Court judgments. Every pipeline shares one design choice that keeps the benchmark honest: a single gold trace per scenario, authored once and shared unchanged across every language variant, so a model cannot be handed an easier question in one language than another a property we verify with a build-time test rather than assert. Each episode is scored on four signals : execution, process, call-matching, and policy : combined as a product, so a fluent response in one turn cannot compensate for a safety failure in another. Before trusting any model's score, we validated the scorer itself against an 8-axis,  55,000-case mutation and golden-replay suite. We report results for twelve open-weight and frontier models, and use two of them for a detailed case study: tracing a low score to a specific, ground-truth-verified mechanism, a model declaring intent to call a tool, then answering as though it already had, producing a fabricated citation in the process and testing a minimal, disclosed prompt intervention against it. The corpus, harness, and scorer are released alongside this paper.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.