AcoustiGym: Physics-Based Acoustic Environment Generation for Robotic Policy Evaluation
Abstract
Spoken instructions reach a robot through a room and a speech-recognition pipeline, either of which can change the goal selected from the instruction. We present AcoustiGym, which exports generated indoor scenes for acoustic simulation, propagates synthesized speech to a microphone array, and evaluates the top-ranked goal selected from the resulting transcript. We compare written instructions, clean synthesized speech, and room-propagated speech to separate the written-to-spoken grounding gap from the additional effect of room acoustics. In one generated bedroom, a predicate resolver grounds of written instructions correctly and after clean speech synthesis and recognition across 50 goals with four phrasings. Across four furnished-mesh room variants, the additional room-only regret ranges from to ; the hard-floor variant has a 95% confidence interval of . In a separate shoebox-room study, an expert-guided room configuration increases room-only regret to , and a one-surface replacement reduces it to zero on 24 evaluation items.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.