acceptodds
Under review as a conference paper at ICLR 2027

Underused Verifier Signals? Refine and Replay for Multi-Constraint Instruction Following

Abstract

Multi-constraint instruction following requires a model to satisfy content, format, length, and lexical requirements jointly within one response. Group-relative reinforcement learning with verifiable rewards offers scalable supervision, but it typically compresses each verifier vector into a scalar reward. This compression obscures which constraints fail and whether equally rewarded responses require different corrections. We find that the discarded verifier structure can refine response-level credit, while successful on-policy siblings provide useful but potentially unsafe token-level supervision. Building on these observations, we introduce Verifier-Informed Refinement and Replay Optimization (VIRRO), which retains the scalar group-relative advantage as a carrier and adds two controlled learning paths.Reliability-calibrated alias refinement uses past-only constraint-family histories to differentiate responses with identical scalar credit while preserving carrier-resolved order and sign under a relative norm cap. Termination-safe peer replay reinforces a verified successful sibling only when termination, length, advantage, and same-group failure conditions admit the route. Across five instruction-tuned backbones and four programmatically checked benchmarks, VIRRO improves the corresponding base checkpoint on all 40 strict metrics and attains the highest local mean in 37 of 40 comparisons. Code is available at this anonymous repository.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.