acceptodds
Under review as a conference paper at ICLR 2027

Same Model, Different Interactions: Hidden History in Multi-Task Learning

Abstract

Multi-task learning routinely characterizes the interaction between tasks from an instantaneous snapshot of the current model, using quantities such as task losses, gradients, or pairwise gradient alignment. Such snapshots describe how tasks relate at a given moment, but treating them as sufficient to determine how those relations evolve under continued training is an assumption, not a guarantee. We show it fails in the strongest possible sense: for any current model, one can construct distinct optimization states that implement it exactly, are indistinguishable under any model-derived instantaneous representation, yet evolve into different future task–response maps. What is missing is therefore not a better observable but the optimization history itself. That decisive history lives in the optimizer's momentum, which carries a hidden, task-attributed record of past training rather than unstructured noise. As a consequence, the same pair of visible tasks can help each other under one hidden history and harm each other under another. More strongly, two indistinguishable states can call for different next tasks to train: we quantify this with an Aliasing Decision Gap and prove that any policy acting on the snapshot alone incurs an irreducible regret, which we confirm across model families and task suites over many seeds. Task interaction is therefore not an intrinsic property of a task pair but a history-conditioned relation carried behind an identical model and exposed only through history.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.