acceptodds
Under review as a conference paper at ICLR 2027

The Perception-Override Reflex: When Speech Language Models Follow Perception Over Instructions

Abstract

Speech language models have become increasingly capable of recognizing paralinguistic information such as gender, emotion, and accent from audio. However, we find that their performance degrades sharply when an affirmative instruction is replaced by its negated counterpart, with models frequently returning the option explicitly excluded by the instruction when that option is supported by the audio. We term this failure mode the perception-override reflex. To characterize it, we introduce Override, a controlled benchmark that evaluates 4,204 audio clips under matched affirmative, negated, and text-control variants, allowing us to control for affirmative-task errors and text-side instruction execution. Across seven speech LLMs, negated instructions induce substantial accuracy drops even after restricting evaluation to items answered correctly under the corresponding affirmative instruction. We then ask what predicts variation in the reflex. Across 56 phrasings, text-side executability consistently predicts susceptibility under audio across six models, but does not fully account for the observed behavior: substantial reflex remains for highly executable instructions, and counterfactual formulations exhibit additional difficulty beyond that captured by text-side executability. These findings motivate on-policy self-distillation from the frozen model conditioned on the attribute in text and a reliably executable canonical phrasing. On Qwen2.5-Omni, this raises aggregate negated-instruction accuracy from 63.7% to 83.6% without degrading affirmative audio performance. Together, our results identify a reproducible failure in combining auditory evidence with negated instruction following, and show that text-side instruction competence provides both a predictor of vulnerability and a useful signal for mitigation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.