RepDebt: The Debt Direction Surgery Cannot Replay
Abstract
Multi-task learning hopes similar tasks reinforce each other; interference is just as documented, and few instruments measure either effect. We built one — and it surfaces a failure mode direction surgery cannot see. SphereBench-54 is a benchmark of 54 datasets over 7 modality families, mixed so edge-modality influence stays measurable. We train 423 models under one protocol: 99 jointly on all 54, dense controls parameter-matched at 1.5M, plus 324 per-dataset specialists. We vary only the input representation of four point-cloud datasets (flat serialization, multi-view rendering, Hilbert-curve ordering); the other fifty are byte-identical. This harmless-looking choice decides the fate of plain shared trunks (one encoder, no routing). Our best such trunk swings from the best main-scan macro-accuracy among parameter-matched jointly trained models, 0.636 under one serialization, to 0.392 under another (paired Wilcoxon, ). Expert-routed architectures (MMoE, PLE, Cross-Stitch) at equal or greater budgets move by at most 1.62 points; no expert cross-serialization comparison survives Bonferroni correction. This is, to our knowledge, the first controlled measurement of interference cascading through a shared encoder into byte-identical untouched tasks (the 14 vision datasets drop from 0.46 to 0.19). Gradient probes of retained checkpoints show why: a norm-dominated rank-one capture leaves the collapsed trunk's joint gradient nearly rank-one (mean pairwise task-gradient cosine +0.83 versus 0.08 elsewhere); routed and latent-bottleneck models keep per-task gradients orthogonal. Follow-up controls close the loop: the frozen readout is a minor tax (4.8 of 24.5 points), direction surgery repays nothing, per-task gradient-norm equalization repays 23.7 — the carrier is the norm-dominated rank-one update — and a vanilla dense trunk without the spherical-iterative core never incurs the debt: it is specific to shared iterative dynamics, not dense sharing. Routing buys robustness, not sample efficiency. How far this generalizes is what the benchmark is for. Code, configurations, and the evaluation protocol are released at https://anonymous.4open.science/r/RepDebt-A449.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.