acceptodds
Under review as a conference paper at ICLR 2027

ActionWhisper: Hijacking What Voice Agents Do, Not Just What They Say

Abstract

Voice agents translate spoken requests into structured tool calls, exposing application state to adversarial audio. An attacker-specified action requires the model to generate a valid call with the intended tool and exact arguments. Revisiting response-level audio attacks reveals that an attack can find a waveform that executes the target action, yet return a failed candidate with lower response loss. Teacher-forced loss evaluates the target sequence under correct prefixes, so it can guide waveform updates without reliably ranking freely generated calls. We propose **ActionWhisper**, which uses this loss for search and execution outcomes for candidate retention. Starting from a bounded speech-difference perturbation, it assigns distinct weights to call-format, tool-name, and argument tokens. A two-phase schedule emphasizes format and tool choice before increasing argument emphasis, keeping the complete call in the objective. At regular intervals, it decodes without target prefixes and evaluates the resulting calls through the parser and executor. Candidate selection prioritizes exact action matching, target-tool invocation, and call validity, using response loss to break ties. Across four Audio LLMs and two agent frameworks on executable cases, ActionWhisper achieves the highest target-tool invocation and precise-action rates in all eight settings under matched perturbation budgets and step caps, raising macro-average precise-action rate to , versus for the strongest baseline, Target-Text CE. Code is available at https://anonymous.4open.science/r/actionwhisper-review/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.