acceptodds
Under review as a conference paper at ICLR 2027

Learning to Reason from Reference Samples Alone: Reinforced Adversarial Learning for Modern LLM Post-Training

Abstract

In many settings where LLM post-training is desirable, high-quality answers are recognizable but difficult to score with explicit rewards. Reinforcement learning with verifiable rewards (RLVR) is limited to tasks where an explicit correctness signal is available. Prior adversarial approaches to training reasoning LLMs either complement verifiable rewards with co-trained critics or require reference reasoning traces. We ask whether a co-trained LLM judge can turn reference answers alone into the sole reward for on-policy post-training of reasoning LLMs. We call this method Reinforced Adversarial Learning (RAL). We first validate RAL on three verifiable tasks (integer multiplication, code-based arithmetic-expression synthesis, and SMILES-to-formula translation), using the verifiers only for evaluation. Judge-only RAL, with sufficiently capable judges and enough training, matches or exceeds direct RLVR. In the first, we ask the model to generate paper abstracts without specifying the target domain in the prompt, but the judge can infer that domain and provide feedback that guides the model toward it. In the second, we generate abstracts for U.S. National Institutes of Health grant proposals conditioned on the grant title. Human scientists and an ensemble of LLM judges find our generations harder to distinguish from real abstracts than those of frontier-LLM baselines. These results suggest a path toward post-training reasoning models in scientific and other high-value domains, where increasingly capable judges infer hard-to-specify reward functions from example answers. %We release our code and datasets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.