FINE-GRAINED DEVICE-DIRECTED SPEECH RECOGNITION: FROM OFFLINE MODELING TO STREAMING INFERENCE
Abstract
A smart speaker in a shared space must select device requests from mixed conversation, adapt to changing functions, and preserve information beyond the transcript as interaction unfolds. We introduce fine-grained device-directed speech recognition (DDSR), which jointly transcribes semantic units and identifies their intents. An offline realization establishes this recognition and selection interface, with contextual verification for unresolved intents. Because selection depends on the device's supported functions and users' expressions, dialogue-based intent management updates the candidate inventory and registers difficult expressions without changing model weights. To retain information that transcription alone omits, the interface associates each semantic unit with speaker, emotion, and acoustic-event information. We then extend the interface to continuous input through fine-grained streaming device-directed speech recognition (FSDR). Intent-delimited unit completion enables incremental selection, while parallel attribute heads provide rich outputs without inserting additional attribute tokens into autoregressive transcription. On English and Chinese evaluations, offline DDSR achieves recording-level equal error rates of 1.25% and 5.46%, compared with 2.74% and 6.26% for a reproduced SELMA DDSD baseline; FSDR achieves 4.60% and 9.82%. Further evaluations assess dynamic intent management, acoustic-event classification, emotion recognition, and speaker attribution, connecting fine-grained content selection with configurable, informative, and incremental outputs for voice assistants.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.