acceptodds
Under review as a conference paper at ICLR 2027

Mergodic: Almost-Lossless Model Merging

Abstract

Fine-tuning a shared pretrained model yields a rapidly growing population of task-specific experts, and model merging has emerged as the standard way to consolidate them into a single deployable artifact without retraining. Existing approaches face a fundamental tension. Static merging methods collapse all experts into one set of weights and are compact, but parameter interference causes accuracy to degrade as the number of tasks grows. Task-conditioned methods, often implemented with routing mechanisms, recover much of this lost accuracy by retaining task-specific signal, but they reintroduce exactly what merging set out to remove: per-task masks, adapters, or expert modules whose storage scales with the number of tasks. We ask whether a single shared representation can instead regenerate any requested expert on demand. We answer in the affirmative with Mergodic, a task-conditioned merging method grounded in a classical result from ergodic theory: by the Kronecker-Weyl equidistribution theorem, the orbit of a one-dimensional irrational rotation is dense in the -dimensional torus, so the cross-task values at any parameter coordinate can be approximated by a single scalar phase. Mergodic discretizes this phase into one integer code per coordinate—shared across all experts and independent of the number of tasks—and reconstructs each expert through a deterministic, task-conditioned rotation map. The result is an almost-lossless merge: experts are recovered with negligible reconstruction error at a fraction of the storage of the original collection. We evaluate Mergodic under both expert-availability regimes: offline merging, where all experts are available simultaneously, and continual merging, where experts arrive sequentially and previous checkpoints are discarded. Across vision and natural language benchmarks, Mergodic attains state-of-the-art performance while remaining highly compact.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.