acceptodds
Under review as a conference paper at ICLR 2027

GRPO-JEPA: Mitigating Semantic Drift in GRPO via Joint Embedding Predictive Architecture

Abstract

Group Relative Policy Optimization (GRPO) is a standard method for RL post-training of large language models, yet its objective stays in the output space: neither the policy gradient nor the token-level KL term observes representation-space dynamics. This blind spot lets semantic drift occur while the loss stays low. We present GRPO-JEPA, which brings the Joint-Embedding Predictive Architecture (JEPA) into LLM RL post-training as explicit representation-space supervision: an advantage-weighted InfoNCE signal turns GRPO's group-sampled completions into multi-view data, propagating scalar reward signals into representation-space alignment. This JEPA supervision signal enters training in two ways: (i) Method A (direct regularization) adds it to the GRPO objective, and (ii) Method B (soft modulation) converts it into a dynamic controller of the KL coefficient. Across four benchmarks (Spider, GSM8K, HellaSwag, and Yelp) and five models from four families (Qwen, OLMo, SmolLM, and Gemma), GRPO-JEPA improves on standard GRPO on all model–dataset pairs: average gains are points for Method A and for Method B. It also surpasses recent GRPO variants (SAPO, BNPO, and DAPO) on every benchmark for Method A and on average for Method B. Representation-space predictive consistency offers a usable JEPA supervision signal for RL post-training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.