acceptodds
Under review as a conference paper at ICLR 2027

Lying in VQA Without Getting Caught: CoT-Obfuscated Backdoors in VLMs

Abstract

Generative (thinking and reasoning) models can justify their outputs in the form of Chain-of-Thought (CoT) transcripts. However, it remains debatable how accurate these justifications are in practice. In this paper, we take this thought one step further and present a neural backdoor that disguises the attack in light of CoT reasoning. The backdoored model is tasked with Visual Question Answering (VQA), but answers wrongly if the input image contains a certain “trigger pattern” and lies about it, meaning the CoT transcript provides a reasonable justification why the answer would be correct nevertheless. Our attack, Chain-Of-Lies, consists of two fundamental training phases, Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO), in which we first implant the backdoor into the base model and then train the model towards stealthiness using Reinforcement Learning (RL). For the latter, we use rewards from a separate neural overseer VLM, specifically instructed to penalize the model if the CoT reveals the backdoor. In our evaluation, we show that our attack bypasses defense mechanisms such as “Chain-of-Scrutiny” and generalizes across multiple VLMs, including Qwen3-VL , Gemma3 , and Ministral-3 . Our results show that an AI model’s self-reasoning cannot be trusted for safety evaluations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.