acceptodds
Under review as a conference paper at ICLR 2027

Depthwise Channel-Asymmetric Attention: Investigating Capacity, Scaling, and Uses

Abstract

Vision transformers embed image patches by summing over input channels, which assumes every channel measures the same quantity at the same point. This fails for instruments whose channels are separate sensors, such as calorimeter detectors, where electromagnetic and hadronic channels are physically distinct. DepthViT is an architecture which attends along the channel axis instead of the token axis. We extend this model for more general image tasks by restoring the resulting loss of spatial communication with sparse hierarchical attention pooling (HAP) blocks and a structure-preserving head. At a matched 22M budget on the HLS4ML Large Hadron Collider jet benchmark, DepthViT2 reaches 75.34% top-1 against 73.98% for a channel-symmetric ViT-Small, winning every class and using fewer FLOPs; on CIFAR-100 it leads 62.76% to 52.92%. The advantage is budget-dependent: across a ladder from 164K to 22M parameters the comparison changes sign, losing at 164K (-1.56 points), tying at 1M, and winning at 5.4M and 22M (+1.69, +1.36). We trace this to parameter allocation, since the channel operator scales quadratically in the capacity knob k while the feed-forward block is linear, and confirm it with a same-scaffold control showing a parameter-matched symmetric operator matches accuracy, so asymmetry is the cheapest configuration rather than the most accurate one. A scale-adjusted positional encoding further lets one trained model accept many input shapes, beating a fixed-crop ViT that discards an average of 44% of each image. Finally, matched FLOPs coexist with a 4x wall-clock cost that a fused Triton kernel removes at 71x on the isolated operator.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.