acceptodds
Under review as a conference paper at ICLR 2027

Societal Hacking: Reward Optimisation Amplifies Societal Loophole Exploitation in LLMs

Abstract

Reinforcement learning (RL) has become the dominant post-training paradigm for large language models (LLMs), and its characteristic failure is reward hacking, where an ill-defined reward function acts as a proxy that captures the measurable form of an intent rather than the intent itself, and optimisation finds the gap between the two. Many codified rules human society runs on are proxies of the same kind. We ask whether reward hacking in LLMs then scales into societal hacking, the unprompted pursuit and exploitation of loopholes in the rules society runs on. To study this safely, we build a set of simulated institutional environments, reverse-engineered from real-world regulations and from documented vulnerability patterns for analysis. Within these environments, reward hacking emerges without any instruction to look for loopholes. The LLM rediscovers the loopholes that later amendments closed, produces strategies that remain formally compliant while defeating regulatory intent, and passes through refusal safeguards that block the same content when it is requested directly. Patching loopholes as they appear does not eliminate the underlying behaviour but merely redirects it towards subtler forms. Collecting in-the-wild feedback for model training therefore demands greater caution and calls for a next-generation post-training paradigm for safely iterating LLMs in real-world environments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.