acceptodds
Under review as a conference paper at ICLR 2027

Collaborative Learning and Reasoning over Model Slices

Abstract

Large language models (LLMs) excel across diverse tasks, benefiting from strong knowledge priors, flexible in-context reasoning, and scalable test-time computation. Despite these strengths, single-model reasoning remains vulnerable to error propagation, as later steps depend on the model's earlier outputs, leaving reasoning blind spots difficult to independently reassess. To address this limitation, we develop *collaborative learning*, encompassing the construction of heterogeneous collaborators and the learning of effective cross-model collaboration. Specifically, *model slicing* replaces a large model with smaller, heterogeneous models under a comparable parameter budget, enabling prior reasoning to be reassessed across distinct model parameterizations. *Pyramid supervision* combines protocol distillation with step-level supervision to refine reasoning and interaction decisions throughout collaboration. We evaluate our approach across four task categories: multi-hop question answering, mathematical reasoning, code generation, and open-ended constrained generation. The trained heterogeneous models achieve comparable or better task performance than the original large model and established multi-agent system and multi-agent reinforcement learning baselines, while achieving higher inference efficiency than the evaluated multi-agent frameworks and exhibiting differentiated collaboration patterns across tasks. These findings show that learned collaboration turns model heterogeneity into complementary reasoning capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.