Shared-Factor Distribution Shift in Multi-Task Learning: Identification and Adaptation
Abstract
Multi-task models serve tasks whose labels arrive at different speeds. Clicks arrive within minutes and purchases only after weeks. After a shift, the slow tasks must be predicted from the fast tasks' early labels. Fine-tuning on those labels improves every task that can be checked, and worsens a slow one. Which part of a slow task's shift do the early labels determine, and what should be done about the rest? We formalize the shift as window-specific gains on a few factors shared by all tasks, and derive what the early labels determine. A slow task's shift is identified if and only if every factor it reads is read by some fast task, and two past windows identify the factors themselves. Because the part that no fast task reads is unknown, an update that reaches into it can lose by any amount; the best update that avoids it is the refit with its unread part removed. FLAG serves this projected update on a frozen multi-task model. Replayed on four datasets and 47 weeks of a question-answering site, it has the lowest slow-task error of eight baselines, and is worse than not adapting in none of 89 served windows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.