AuralLoop: Language-Guided Encoder Re-entry for Audio Question Answering
Abstract
One-pass audio-language models complete acoustic encoding before language-side inference, so intermediate language states cannot guide further encoder computation. We introduce AuralLoop, a parameter-efficient architecture that makes acoustic encoding responsive to evolving language context. A shared low-rank cell conditions a cached acoustic anchor on an intermediate language state and reruns the original frozen encoder tail. Incremental embedding differences are injected into existing audio-token positions at later language layers. Two re-entry steps share the anchor and encoder weights, with the second conditioned on a language state that has incorporated the first revision. Instantiated on Qwen2.5-Omni-7B, AuralLoop uses 3.20M trainable parameters while keeping pretrained backbone weights frozen. It requires no additional encoder, generated control tokens, or chain-of-thought supervision during adaptation. Evaluations on ADQA, VoxParadox, MMSU, and MUGEN compare re-entry with direct acoustic adaptation, static feature access, and additional training computation. On VoxParadox, AuralLoop improves mean accuracy by 1.10 percentage points over a single re-entry at a later language layer and by 0.97 points over joint audio-language LoRA under a controlled training-work budget. Same-depth interventions suggest task-dependent benefits from updated language conditioning, while capability analyses show the largest gains in paralinguistic perception. Paired error and latency analyses characterize the accompanying answer regressions and prefill cost. These results support language-guided encoder re-entry as a targeted approach to acoustic question answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.