Underused Verifier Signals? Refine and Replay for Multi-Constraint Instruction Following
Abstract
Multi-constraint instruction following requires a model to satisfy content, format, length, and lexical requirements jointly within one response. Group-relative reinforcement learning with verifiable rewards offers scalable supervision, but it typically compresses each verifier vector into a scalar reward. This compression obscures which constraints fail and whether equally rewarded responses require different corrections. We find that the discarded verifier structure can refine response-level credit, while successful on-policy siblings provide useful but potentially unsafe token-level supervision. Building on these observations, we introduce Verifier-Informed Refinement and Replay Optimization (VIRRO), which retains the scalar group-relative advantage as a carrier and adds two controlled learning paths.Reliability-calibrated alias refinement uses past-only constraint-family histories to differentiate responses with identical scalar credit while preserving carrier-resolved order and sign under a relative norm cap. Termination-safe peer replay reinforces a verified successful sibling only when termination, length, advantage, and same-group failure conditions admit the route. Across five instruction-tuned backbones and four programmatically checked benchmarks, VIRRO improves the corresponding base checkpoint on all 40 strict metrics and attains the highest local mean in 37 of 40 comparisons. Code is available at this anonymous repository.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.