acceptodds
Under review as a conference paper at ICLR 2027

Unhack: Robust RLVR Under Unreliable Specifications

Abstract

Reinforcement Learning from Verifiable Rewards (RLVR) improves models’ ability to solve tasks in isolation. In deployment, however, models interact with users, tools, and other models, which introduce additional context that may be wrong or outdated. We find that RLVR-trained models often follow this context even when it conflicts with the task objective, including on problems they solve correctly without it. We study this behaviour in three settings: mathematical reasoning with answer hints, code generation with tests, and agentic tool use with action plans. We refer to it as specification anchoring, examine how RLVR training affects it, and introduce UNHACK, an RLVR training method designed to mitigate it. Untrained models already exhibit the behaviour; standard RLVR does not remove it in any domain we test and sometimes amplifies it. Across six model configurations, UNHACK substantially reduces compliance with corrupted specifications in all three settings while matching or improving standard task performance. The improvement generalises to corruption families outside the training support, while models continue to benefit from correct specifications. Common test-time mitigations, including prompt warnings and self-verification, do not reliably prevent this behaviour.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.