acceptodds
Under review as a conference paper at ICLR 2027

CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) and Test-Time Scaling (TTS) have advanced LLM code generation through executable verification, yet both hinge on Ground-Truth Unit Tests (GT UTs), which are expensive to annotate at scale. GT-free methods instead rely on self-generated UTs to verify and refine codes, but these UTs are often noisy, and so are the codes they check. Treating either side as ground truth to improve the other thus only propagates its errors. The key challenge is therefore: with no external reference, how can each side extract usable signals from the other, turning error propagation into mutual error reduction? We present CoSPlay, a GT-free framework that co-evolves codes and UTs through cooperative self-play at inference time. Although neither side is reliable alone, we show theoretically and empirically that pass counts on the Code-UT execution matrix correlate positively with correctness and can thus, to some extent, tell trustworthy codes and UTs from unreliable ones. Building on this signal, CoSPlay first derives discriminative UT ideas from the failure modes of diverse solution ideas. It then alternately updates the two pools, keeping the codes and UTs with higher support counts and using them to fix or regenerate the rest. Finally, such a signal indicates the model does not lean toward any particular wrong behaviors in expectation, so the correct behavior dominates the top codes, whose ties CoSPlay breaks by unsupervised clustering of output signatures on a few random inputs. On four challenging benchmarks, CoSPlay improves the average Best-of-N (BoN) accuracy of Qwen2.5-7B-Instruct from 22.1% to 33.2% and its UT accuracy from 14.6% to 78.3%. Without any GT UTs, this matches or surpasses CURE-7B, an RLVR model on the same base, trained on 4.5k samples with GT-UT rewards for both code and UT generation, and CoSPlay further improves CURE-7B itself by 5.7% BoN. CoSPlay also generalizes across diverse backbones and achieves higher average accuracy than GT-free TTS baselines under comparable token budgets. By letting codes and UTs improve each other without any GT, CoSPlay offers not only a scalable inference strategy but also a new perspective on recursive self-improvement for verifiable problems without external supervision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.