acceptodds
Under review as a conference paper at ICLR 2027

Aggregate, Gate, and Guide: Towards Efficient, Balanced, and Sharp Multi-Task Dense Scene Understanding

Abstract

Multi-task dense scene understanding jointly solves multiple pixel-level tasks with a shared image encoder, aiming to exploit task correlations while maintaining computational efficiency. However, existing methods often perform cross-task interactions at multiple feature scales, introducing considerable computational overhead. Moreover, they typically exchange task information without explicitly preserving task-specific patterns, which may compromise balanced performance. Meanwhile, their upsampling operations are commonly performed independently for each task, overlooking complementary task cues that can facilitate the recovery of fine structures. To address these limitations, we propose AGG-MTL, a multi-task dense prediction framework that advances the efficiency–balance–sharpness trade-off. First, we aggregate multi-scale task features into compact representations and perform cross-task interaction only at a single scale, reducing repeated computation across resolutions. Second, an adaptive cross-task interaction transformer employs soft-to-hard gating to selectively incorporate cross-task cues while preserving critical task-specific patterns. Third, a multi-task guided upsampling module predicts task-specific interpolation offsets from complementary multi-task features, extending cross-task collaboration to fine-structure recovery. Experiments on NYUDv2 and PASCAL-Context demonstrate a favorable trade-off among computational efficiency, balanced multi-task performance, and prediction sharpness. On NYUDv2, AGG-MTL improves by +3.54 points over the most efficient prior method, reduces FLOPs by 15.7% relative to the previous best-balanced method, and achieves the best sharpness scores across all tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.