AMIYC: Add Me if You Can Conditional Adaptation Headroom and Recovery in Evolving MMoEs
Abstract
Multi-task mixture-of-experts (MMoE) models are commonly evaluated through routing behavior and end-to-end task quality, but neither establishes whether a trained state has exhausted the task-conditioned improvement available within its learned representation. We call this measurable opportunity conditional adaptation headroom. We introduce AMIYC, a framework that separates routing exposure from the task-visible expert state and distinguishes fresh adaptation opportunity, compatibility of earlier interventions, and recovery from a finite action bank. Across Amazon Reviews (Beauty, Grocery, Sports, and Toys), KuaiRand-Pure, and a production e-commerce ranking MMoE, bounded interventions reveal positive headroom on the tested cohorts, including near validation-selected checkpoints and when the same intervention coordinates participate in ordinary training. Alternative adapters show that the reference diagonal is not unique. As representations evolve, earlier operators may retain value, become harmful, or recover usefulness, while tested transport corrections do not consistently improve replay. Targeted Amazon studies further reveal reversals in checkpoint ordering after common-anchor fitting and associate a tested coupled adaptation schedule’s deficit predominantly with its base trajectory. Realizing this opportunity remains constrained: most final raw gain in the tested continuation is inherited from the historical anchor, affine calibration substantially reduces the apparent loss advantage, and useful ranking improvements remain selective. Finite-bank availability, action selection, and secondary-quality constraints further limit recovery. Fresh-seed confirmation reuses previously examined data. Together, these findings distinguish bounded, state-relative opportunity from compatibility and operational recoverability, without implying adaptation saturation or a universal remedy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.