Robust to Noise, Broken by Turns: Benchmarking Speech-Driven Omni Agents in the Wild
Abstract
Speech is a natural interface for agents, yet most agent benchmarks assume typed instructions. We introduce OmniAgentBench, a benchmark for evaluating speech-driven multimodal planning across Schedule, GUI, and Embodied tasks, with matched text and audio inputs and controlled wild conditions. Native omni models perform comparably to ASR+LLM cascades (automatic speech recognition followed by an LLM) on clean speech, but they collapse when the same constraints are distributed across dialogue turns, with Qwen-Omni-3 dropping from 62.0 to 18.8 action matching score (AMS) on GUI navigation. Controlled experiments trace this degradation to the plan-every-turn protocol, not to speech perception or ASR. Because the model must respond after each turn, it commits early to an incomplete plan. Presenting the same turns as text produces a similar collapse, while suppressing intermediate responses largely restores performance. Training-free prompts that make the model wait for the last turn, or restate every constraint before planning, recover most of the loss without extra model calls, and the collapse replicates across four native omni models from three labs. We also recorded 116 human speakers reading the same instructions. On clean speech, human recordings score about as well as TTS on average, and the multi-turn collapse reproduces on real voices. OmniAgentBench contains 11,700 audio cases from 900 base instances and provides a reproducible protocol for studying speech-grounded multimodal planning under realistic instruction variation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.