acceptodds
Under review as a conference paper at ICLR 2027

The Scar of RL: Membership Inference Attack for Reasoning Models

Abstract

Reinforcement learning (RL) is a critical post-training stage of frontier reasoning models. However, the prompts used during RL training are rarely disclosed, making it difficult to determine whether benchmark questions used to evaluate post-trained models were included during RL training, and potentially biasing downstream evaluations. Existing membership inference attacks, developed primarily for pretraining and supervised fine-tuning, typically exploit the reduced loss of an observed training sequence. However, this principle does not transfer directly to RL where there are no observed sequences. Instead, RL optimizes for rewards given to sampled outputs and reshapes the distribution over model-generated solution trajectories. We show that this change leaves a multi-dimensional signature in the model’s own outputs as prompts seen during RL tend to induce higher confidence in sampled continuations, more similar solution trajectories and final answers, more saturated verifier outcomes, and less variation across independent rollouts. However, the magnitude of these effects varies substantially across model–dataset pairs, making any single statistic unreliable. In response, we introduce SCAR, which defines 15 features capturing complementary signatures of RL exposure, and combines them through label-free multi-axis aggregation, avoiding reliance on any single signature. Across eleven unsupervised aggregation rules, combining the four axes (Self-sharpening, Collapse, Answer saturation, and Rollout-sharpening) consistently improves macro-average performance over individual signatures. Our strongest aggregation reaches mean ROC-AUC, compared with for the strongest single-statistic baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.