acceptodds
Under review as a conference paper at ICLR 2027

The 48 Laws of LLM Power: Benchmarking and Mitigating Machiavellian Agent Advice

Abstract

Large language models (LLMs) increasingly act as strategic advisers in social and professional settings, creating a safety risk poorly captured by benchmarks focused on direct model-to-user manipulation. We introduce 48LawsBench, which operationalizes Robert Greene’s The 48 Laws of Power – a distillation of three millennia of power seeking – to systematically evaluate whether LLMs act as harmful Machiavellian advisers. We test these laws at different levels of explicitness – from direct requests for unethical tactics to ordinary goals where such tactics are only implicit – distinguishing compliance from cases where models themselves infer harmful strategies. We evaluate 8 commercial and 8 open-weight LLMs and uncover striking differences: without safeguards, the share of responses using high-risk tactics ranges from ∼10% to ∼67% across models. For example, Gemini often recommends concealment, credit appropriation, dependency, and adversarial neutralization, whereas Grok does so substantially less often. Among several recurring cross-model failures, Law 15, “Crush Your Enemy Totally,” is especially persistent despite being judged highly unethical in isolation. Simple prompt safeguards reduce such behavior, but sometimes at the expense of advice utility. We thus introduce GreeneGuard, an infer–audit–rewrite safeguard that jointly targets safety and utility. It further reduces harmful strategic advice while preserving or improving usefulness, addressing residual failures left by simple prompting. Our results show that reducing harmful strategic advice need not come at the expense of effective goal-directed assistance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.