acceptodds
Under review as a conference paper at ICLR 2027

When Should a Spoken Agent Act? A Benchmark for Risk-Aware Function Calling

Abstract

Spoken agents can send messages, control devices, and initiate payments from voice commands. Automatic speech recognition (ASR) uncertainty about a recipient or amount therefore creates an authorization problem: should the agent act, ask one question, or stop? Word error rate and function-call accuracy do not resolve that decision. We introduce Risk-Aware Spoken Function Calling (RASFC), a controlled rendered-speech benchmark that pairs structured calls with risk tiers, argument-support labels, and rule-derived next actions: execute, ask a targeted clarification, or abstain. The current benchmark interface gives its verifier the intended action's risk tier at evaluation time; this oracle context is distinct from risk inferred from speech or a predicted call. In a replayable offline diagnostic on 6,000 distinct audio/ASR observations, reference-WER plus ASR confidence misses 1,147 dialogue-required examples under a rule built from gold argument support, clarification targets, risk, and confidence, whereas a gold argument-support diagnostic misses 389. These signal comparisons measure consistency with the constructed rule. RASFC makes the information flow explicit for evaluating when uncertain speech warrants action, targeted repair, or abstention; it is a controlled benchmark interface, not an end-to-end human-speech deployment evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.