QGA-TTT: Query-Guided Aggregation via Test-Time Training for Medical Image Segmentation
Abstract
Medical image segmentation asks for a label at every position of a scene that holds many structures at once, and which class a region belongs to is decided in large part by what else the image contains. Every position of one image therefore has to be labelled against one and the same account of that image. Convolution reaches only as far as its receptive field, and pairwise attention lets positions compare with one another without producing a single shared description, so neither forms such an account. Test-time training layers fit a small set of fast weights inside the forward pass, so the map they apply is recomputed for every input. We use this property to build one shared account per image. At each of the deepest levels of a segmentation network, the level’s positions together with a small set of learnable queries shared across every image are fitted, in a single forward pass, into one affine map, and every position reads the global information back through that same map. The shared queries enter every fitting, so decoder weights shared across images can read what the map returns. The queries are kept across two passes so that the second reads an account already placed in the coordinates of the current image, and the map is estimated twice by fittings that share no parameters. Across nine public datasets spanning endoscopy, ultrasound, optical coherence tomography and magnetic resonance imaging, the method reaches the best Dice on every one, by up to 2.81 points over the strongest comparison method. Ablations show each of the three constructions to be necessary, and both the number of fittings and the capacity of the account are interior optima.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.