Cross-Family Exchange Under Matched Call Budgets: Sampling, Revision, and Tie Resolution
Abstract
Does cross-family exchange improve a language-model panel beyond independent sampling or revisiting its own responses? We compare Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct on 200 GSM8K questions and a 100-question rational-answer MATH subset. For each question, 16 saved generations support four-call conditions that share initial answers, model composition, output ceilings, and aggregation. Both proposer orders are evaluated. Averaging the two orders within each question, exchange trails independent voting by 3.0 percentage points on GSM8K (descriptive 95% paired-bootstrap interval ) and 1.5 points on MATH (). Relative to own-response revision, the differences are and points; neither interval excludes zero. Exchange adds 603 and 1,405 input tokens per question over independent voting. A post hoc order-invariant tie sensitivity reduces the MATH advantage over own-response revision to 1.1 points: answer order affects the fixed vote even without changing generated content. Shared valid initial agreement also fixes the final vote regardless of subsequent answers. These results identify restrictions of the decision procedure, not a general benefit of exchange or cognitive specialization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.