Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models
Abstract
Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly treating all attacks as equally costly. In practice, the computational expense of different attack strategies can vary by orders of magnitude. Consequently, ASR at a fixed budget can obscure the true effort required to jailbreak a model, thereby making it hard to determine whether an attack’s cost justifies its payoff to the attacker. We propose a compute-aware evaluation framework based on computational pressure, measured in cumulative floating-point operations (FLOPs), as a proxy for adversarial effort. We introduce risk-compute curves, which map compute budgets to attack risk, and derive two summary metrics measuring the compute needed to reach a risk threshold and the risk attained per unit of compute. Across five model families and four stages of training and alignment, we evaluate static and adaptive attacks, including template-based attacks, iterative prompt refinement, gradient-based optimization, and reinforcement learning (RL), on two jailbreak-robustness benchmarks. We find that: (1) alignment training has non-monotonic effects on compute-space robustness; (2) scaling model size reduces attack success, with a more pronounced effect on gradient-based attacks than on template-based attacks; (3) surrogate-optimized gradient attacks transfer without target white-box access, at low per-query cost but limited observed risk; (4) compute cost varies by up to across harm categories within a single model; and (5) safety-aligned RL increases aggregate attack cost while leaving some categories disproportionately accessible. We release our framework to enable compute-aware risk assessment and evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.