acceptodds
Under review as a conference paper at ICLR 2027

Crescendo-Live: Multi-Turn Jailbreaks of Full-Duplex Speech Models in Live Calls and Their Defense with Daimonion

Abstract

Full-duplex spoken dialogue models listen and speak at the same time, aiming at seamless, natural voice conversation. The latency constraint limits the model's ability to reason about what it hears, which leaves it exposed to multi-turn jailbreaks such as Crescendo: the attacker opens with a benign question and escalates gradually, building each request on the model's previous response, so the harmful goal is spread across turns and becomes more implicit. Yet the safety of speech models is still evaluated on single harmful requests or pre-scripted dialogues. We present crescendo-Live, which evaluates full-duplex models in live multi-turn conversation: an LLM-driven user simulator adapts every turn to the model's response and can interrupt the model while it is speaking. Multi-turn escalation jailbreaks every model we evaluate, even those that are robust to direct requests, and attack success grows with the number of turns. Our ablations also find that talking over the model leads to an even higher attack success rate, raising a safety concern unique to full-duplex models. Aligning the speech model itself against such attacks is hard: one hidden goal can be delivered in many different ways, and full-duplex training data is expensive to synthesize. We instead propose Daimonion, which pairs a full-duplex model with an asynchronous LLM safety monitor. The monitor follows the whole conversation, and when it detects a harmful intent, it injects a text instruction into the frontend full-duplex model, which is fine-tuned to turn such instructions into a spoken refusal while keeping its natural conversational ability. Daimonion reduces attack success to below 3% at every conversation length while retaining the model's original turn-taking ability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.