Jailbreaking Large Language Models via Rebuttal Attack
Abstract
Despite recent progress in safety alignment, mainstream LLMs remain vulnerable in rebuttal-style tasks, often prioritizing task completion over safety and producing harmful content. This paper presents Rebuttal Attack, a novel, single-turn, black-box jailbreak method. It rewrites malicious queries as refutable knowledge or viewpoint statements, reliably eliciting unsafe responses with a single prompt without internal model access. In the knowledge setting, we embed factual errors in malicious queries to trigger the model’s error-correction behavior, causing it to produce restricted content while correcting the inaccuracies. In the viewpoint setting, we craft dangerous debate topics that push the model to defend risky positions and reveal harmful details during argumentation. Token-level likelihood analysis and extensive experiments show that Rebuttal Attack substantially outperforms existing single-turn jailbreak baselines on major commercial and open-source LLMs, including GPT-5.1 and LLaMA-4, and generalizes well across models. For reproducibility, our anonymized code is available at https://anonymous.4open.science/r/RebuttalAttack-4745.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.