Residual Aggregation as Nonlinear Scalarization: Exact Gradient Reallocation and Shared-Encoder Transmission
Abstract
Block-structured objectives can encode the same residual evidence yet allocate gradient mass differently depending on how block losses are scalarized. Let block residuals be and losses . We compare linear scalarization with weighted power scalarization , where . We derive an exact pairwise log-odds law showing that native-to-composed allocation distortion separates into relative weight mismatch and times the log-residual ratio. The law yields matching conditions, exact worst-case distortion and inverse rigidity, exponent identification, and state sensitivity. Power-mean and Jacobian-conditioned Gram analyses connect the same scalarization change to candidate rankings and shared-parameter updates. Common-state audits on Energy, Parkinsons, and CelebA-40 quantify the predicted reallocation and its transmission through shared encoders. In paired 200-epoch regression runs, native consistently lowers task-1 error, whereas composed lowers task-2 and worst-task error on both datasets. Macro-average ordering differs across datasets, and the Parkinsons contrast largely attenuates under validation-selected stopping. Together, theory and experiments provide a quantitative account of when scalarization geometry produces persistent, task-selective optimization effects.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.