acceptodds
Under review as a conference paper at ICLR 2027

Read Together, Write Apart: Decoupling Rollout Access from Update Ownership in Multi-Role LLM Training

Abstract

When one language model plays solver, verifier, and reviser, “parameter sharing” conflates two choices: which adapters a role generates from, and which adapter receives its gradient. A shared adapter couples reads and writes; role isolation severs both. We introduce Dual Routing, which makes the two routes explicit: the verifier and reviser generate from a common adapter pool, while each role's update trains only the adapter it owns. With seven independently trained Llama-3.1-8B checkpoints per method, the downstream workflows of shared training and of role isolation erase accuracy from their own first answers in every replicate, whereas Dual Routing's does not, and Dual Routing beats capacity-matched role isolation by 26.7 points in every replicate. A preregistered route factorial shows the gain is learned: read exactly as isolated adapters are read, Dual-trained adapters still win by 30.0 points. Further controls identify the mechanism: training downstream roles on reads at the merged pool's reduced scale is what keeps the workflow from collapsing. Of all trained variants, Dual Routing's reviser is the only one that improves on the untrained model both at repairing wrong answers and at preserving correct ones. Collaborative reads with role-owned writes are a simple design point for training multi-role LLM workflows that adds no forward cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.