acceptodds
Under review as a conference paper at ICLR 2027

Reference-KL Hijacking: Reference Poisoning Attacks in Distributed LLM Reinforcement Learning Frameworks

Abstract

With the growing size of Large Language Model (LLM), distributed LLM reinforcement learning (RL) frameworks have been widely adopted to deploy different roles across multiple remote nodes to alleviate the computational burden of a single server node. However, existing distributed LLM RL frameworks assume that the communication channel and remote nodes are trustworthy, which introduces an opportunity to manipulate remote feedback for adversarial optimization guidance. In this work, we propose Reference-KL Hijacking (RKH), an indirect safety misalignment attack that compromises remote reference feedback used to construct the KL regularization, which adversarially guides policy optimization towards unsafe generation behaviors. Specifically, RKH contains two independent methods, including Feedback Modification (RKH-FM) and Model Replacement (RKH-MR), which target the communication channel and the deployed reference model, respectively. In order to construct the malicious feedback for misalignment, RKH-FM leverages sentence-level safety judgment of generations and predefined values to modify feedback probabilities; and RKH-MR replaces the original benign reference model with an easily obtainable misaligned model, which produces context-dependent feedback probabilities. Extensive experiments on various LLM families and datasets demonstrate that the proposed RKH can effectively impair safety alignment with the preservation of general utility, where RKH-FM induces more severe safety degradation in certain cases and RKH-MR achieves more consistent effectiveness. These results reveal a novel safety risk introduced by the implicit trust assumption to remote reference feedback and the importance of ensuring the integrity and trustworthiness of remote feedback in distributed LLM RL frameworks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.