acceptodds
Under review as a conference paper at ICLR 2027

Less Fusion, Same Action: Revisiting Redundant Connectivity in Mixture-of-Transformer based Robot Policy Models

Abstract

Mixture-of-Transformers (MoT) robot policies typically maintain dense expert connectivity across network depth, yet it remains unclear how much of this connectivity is actually necessary for downstream adaptation. We conduct a systematic empirical study of expert connectivity redundancy by sparsifying pretrained MoT policies during fine-tuning across multiple architectures and robot manipulation benchmarks. We find that substantial portions of dense expert connectivity can be removed while largely preserving downstream performance, and that different sparse masks with matched budgets achieve similar results, suggesting that no particular subset of expert connections is uniquely indispensable. Motivated by this flexibility, we introduce an early-half fusion pattern that concentrates cross-expert connectivity in the first half of the network and applies structured pruning to the remaining layers. This design maintains comparable task success while enabling practical deployment simplifications, including truncation of unused conditioning layers and reduction of historical KV-cache storage in autoregressive policies. Deployment measurements further show reductions in loaded parameters, GPU memory, and per-chunk latency. Together, our results suggest that dense expert connectivity in pretrained MoT robot policies is substantially redundant, and that structured pruning provides a simple path toward more efficient deployment without training a new model from scratch. Our code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.