Text-Assisted Unsupervised Domain Adaptation for Image Regression through Semantic Re-encoding
Abstract
Unsupervised domain adaptation for image regression transfers continuous prediction rules without target labels. Existing approaches align visual representations, but feature alignment alone need not preserve the numerical relationships required for regression. Pretrained text encoders offer an alternative representation of numerical attributes, yet target text must be constructed from uncertain predictions. We first analyze the reliability tradeoff of this representation. Deterministic target text adds no information conditional on its complete visual input, but can benefit a restricted regressor through re-encoding. A target absolute-risk bound separates source fitting, domain discrepancy, predictor compatibility, and text-construction error; squared-loss analysis further characterizes branch agreement and conditional error propagation. These results motivate retaining visual information, attending to pseudo-label reliability, and jointly supervising modality-specific and fusion predictions. Guided by these insights, we propose TCC, which encodes source labels and target pseudo-labels, trains three regression branches with a residual-informed objective, and updates the target text by fusing predictions with the preceding text. Experiments on dSprites, MPI3D, and BIWI evaluate the framework through baseline comparisons, branch ablations, and representation and parameter diagnostics. Our source code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.