acceptodds
Under review as a conference paper at ICLR 2027

EMIT: Measuring the Reliability of Per-Prompt Decisions in RLVR Curricula

Abstract

Per-prompt controllers in reinforcement learning with verifiable rewards transform small groups of verifier outcomes into consequential training actions, yet the reproducibility and semantic fidelity of those decisions are rarely characterized. We introduce EMIT, a reliability-aware evaluation framework that treats adaptive curriculum control as a measurable decision process. EMIT formalizes exact-action reproducibility, connects difficulty concentration and decision boundaries to rollout requirements, and combines these predictions with source-traced implementation and verifier audits. Applied to frozen real-server rollouts, the framework reveals how sampling resolution and answer-format semantics govern controller behavior. A completion-token-matched multi-run study further shows substantial observed gains for both the released adaptive controller and a carefully tuned static curriculum over the no-curriculum reference. The adaptive controller attains the highest observed primary-task mean while closely tracking the tuned static curriculum across the broader suite. By making intermediate curriculum decisions auditable, EMIT provides a principled basis for reporting, sampling, pooling, and reliability-aware controller design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.