acceptodds
Under review as a conference paper at ICLR 2027

CROSS-TASK GAINS WITHOUT AUXILIARY LABELS: ARCHITECTURE-PRESERVING ABLATIONS UNDER DISJOINT SUPERVISION

Abstract

When two tasks are labelled on different datasets and no example carries both, coupling them through a shared module often improves the target task, a gain credited to the auxiliary labels. But the module also changes the architecture available to the target task, so comparing coupled with separate heads cannot separate supervision transfer from architectural benefit, and none of the fourteen multi-objective methods and systems we audit reports a control that can. We resolve this attribution problem with architecture-preserving label ablations, which need no new model or data: the coupling module and its readout are kept while the auxiliary labels are removed, the module is frozen at random initialization, or the auxiliary task is dropped with the readout kept. A gain that survives without the labels is architectural by construction, and the observed gain splits exactly into architectural and label parts; an exact single-task tie certifies the split. Our testbed is vision-only deepfake localization, where spatial masks and temporal boundaries come from separate corpora, with frozen features and cross-task attention. The control overturns the default reading. Attention improves official [email protected] at every seed, by at ten-fold budget, yet of the pilot gain survives with patch access but no masks (and most of it on a second encoder and dataset pair). At ten-fold, attention without masks matches or exceeds full attention at every seed on the main encoder, and even with the spatial task dropped and the readout kept, most of the gain remains at both budgets (both claims post hoc). A parameter-matched module with masks but no patch access reaches half of the pilot gain; the tested output-space consistency term, a loss through which mask labels do reach the temporal head, gives no consistent advantage. The gain shrinks with budget, is not reliably retained under fine-tuning, and does not predict better localization of unseen manipulation types. A cross-task gain under disjoint supervision should pass this control before it is credited to labels; we release it with scoring code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.