acceptodds
Under review as a conference paper at ICLR 2027

Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It

Abstract

Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy such that its samples increasingly make the same mistakes. With every method drawing exactly completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test (B-B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to points. The cause is concentration, not RLVR itself. We split the same data and training budget across LoRA adapters, each trained on its own random disjoint shard, and call the result an . Thickets out-vote the fully trained adapter in all (model, ) settings, and for they stay within points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small . For , thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from to votes, the thicket's lead over the fully trained adapter widens from to points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.