acceptodds
Under review as a conference paper at ICLR 2027

PoliBench Beyond Refusal: Evaluating LLM Safeguards against Authoritarian Assistance

Abstract

A distinctive safety risk of authoritarian assistance requests is the prompts can be framed as routine governance, compliance, or technical-administration tasks. Recent work has begun to evaluate the sociopolitical harms of large language models, but refusal alone does not establish whether a model will continue to withhold actionable assistance for authoritarian requests. We introduce PoliBench, a benchmark comprising 144 authoritarian assistance goals across 12 operational subcategories, to evaluate 10 target models through three controlled comparisons. Our evaluations reveal that authoritarian assistance requests trigger explicit refusal 20.72% less often than traditional harmful requests. Chinese commercial models explicitly refuse more often than U.S. commercial models, yet their initially refused goals are more often fully achieved under the Tree of Attacks with Pruning (TAP). This strict-yet-porous guardrail pattern raises the threshold for obtaining assistance without establishing a reliable capability boundary. Moreover, for the same goal, models provide assistance more readily when the requester is presented as a non-state institution. Presenting the requester as state-affiliated increases the explicit-refusal rate by 5.70% and decreases the attack-success rate by 5.28%. PoliBench thus shows that guardrails for authoritarian assistance do not form a stable refusal boundary, but a conditional form of control that varies with request content, adaptive reformulation, and requester identity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.