acceptodds
Under review as a conference paper at ICLR 2027

Jailbreaking Large Language Models via Rebuttal Attack

Abstract

Despite recent progress in safety alignment, mainstream LLMs remain vulnerable in rebuttal-style tasks, often prioritizing task completion over safety and producing harmful content. This paper presents Rebuttal Attack, a novel, single-turn, black-box jailbreak method. It rewrites malicious queries as refutable knowledge or viewpoint statements, reliably eliciting unsafe responses with a single prompt without internal model access. In the knowledge setting, we embed factual errors in malicious queries to trigger the model’s error-correction behavior, causing it to produce restricted content while correcting the inaccuracies. In the viewpoint setting, we craft dangerous debate topics that push the model to defend risky positions and reveal harmful details during argumentation. Token-level likelihood analysis and extensive experiments show that Rebuttal Attack substantially outperforms existing single-turn jailbreak baselines on major commercial and open-source LLMs, including GPT-5.1 and LLaMA-4, and generalizes well across models. For reproducibility, our anonymized code is available at https://anonymous.4open.science/r/RebuttalAttack-4745.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.