acceptodds
Under review as a conference paper at ICLR 2027

Learning at Different Speeds: Equilibrium Routing for AI-Driven Science

Abstract

AI for Science increasingly relies on single models to predict multiple observables of the same physical system. However, shared physics does not guarantee synchronized learning: one target may reach its best validation stage early and overfit under continued training, while another still requires optimization. We term this failure mode asynchronous convergence. We introduce Task-Aware Equilibrium Routing Mixture-of-Experts (TAER-MoE), a convergence-aware MoE framework that routes by generalization. TAER-MoE does not treat large loss or high expert affinity as sufficient evidence for assigning more capacity. Instead, it estimates task-level learning potential from validation descent and route-level overfitting from expert-task generalization gaps. A spectral bias–variance perspective links asynchronous convergence to task-dependent physical spectra and different optimal stopping times. Guided by this perspective, TAER-MoE uses an entropy-regularized equilibrium router to suppress expert-task routes with large train–validation gaps and reallocate training pressure toward routes where additional optimization remains useful. Across tropical cyclone forecasting, QM9 molecular property prediction, and a transverse-field Ising model benchmark, TAER-MoE improves joint test prediction over single-task, shared, MoE-based, gradient-based, and ensemble baselines. These results provide empirical evidence that validation-aware routing can improve scientific multi-task learning when coupled observables converge asynchronously.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.