Camera Positional Embedding for Test-Time Training
Abstract
To understand the world in 3D from many images, a model must take into account the cameras that captured them. Camera embeddings for attention do so by relating two tokens only through the relative geometry of their cameras, but they also inherit the quadratic cost of attention, which limits the number of input views. Test-time training (TTT), which stores past tokens in a small network trained at test time, offers an efficient alternative with linear cost. However, how to embed cameras into TTT has not been studied, and camera embeddings designed for attention do not transfer directly. We present CaPET (Camera Positional Embedding for TTT), a camera embedding that acts inside the TTT layer. We build it by asking where a camera embedding should be placed in a TTT layer and what form it should take. For the placement, we show that a TTT layer holds implicit scores inside its fast weights, which let a camera embedding relate two views through their relative camera. For the form, we find that a TTT layer prefers orthogonal embeddings, since they preserve the norms of its keys and queries and so keep training at test time stable. On novel view synthesis across diverse datasets, CaPET improves over the TTT model without camera embedding by up to 1.80 dB PSNR and outperforms camera embeddings for attention, at the linear cost of TTT.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.