acceptodds
Under review as a conference paper at ICLR 2027

Verb-Centric Counterfactual Logic Attacks on Large Language Model Reasoning

Abstract

Recent research indicates that large language models (LLMs) can achieve efficient reasoning by leveraging logical shortcuts, potentially reducing reliance on explicit chain-of-thought processes. However, such shortcuts may introduce vulnerabilities, particularly when reasoning is verb-centric and lacks rigorous logical grounding. In this work, we study counterfactual logic attacks that perturb relation verbs in shortcut triplets. We first show that LLMs exhibit boundary-sensitive preferences over factual, counterfactual, and unrelated verb clusters. To support this analysis, we construct Verb-LK, a curated benchmark of verb-driven logical implications for evaluating model stability under counterfactual logic attacks. We further propose Adaptive Logic Attack (ALA), which uses semantic uncertainty to select logically aligned or contradictory verbs within concept clusters and introduces a contrastive adaptation loss to reconfigure verb-cluster decision boundaries while accounting for potential concept distortion. Experiments show that ALA produces a strong targeted shift in relation-level preferences, generalizes to unseen implications, shows corresponding effects in additional Qwen3-8B experiments, obtains encouraging results in an exploratory anti-stereotype evaluation, and causes only a modest change on MixEval as a broad capability check.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.