acceptodds
Under review as a conference paper at ICLR 2027

TOKAFORMER: ROUTING CO-LOCATED MODALITIES AS TYPED TOKENS BEFORE COMPRESSION

Abstract

Multimodal models often fuse the inputs that share a position by embedding each with one shared encoder and adding the results. This is cheap, but when track identity enters only through additive offsets it erases the pairing between each value and the input it came from, and tasks that must report what one particular input said cannot recover it. We asked when this loss matters and how to avoid it at reasonable cost. For such a shared additive encoding we prove that every reassignment of values among inputs yields the same fused vector, so no downstream model can retrieve a named input's value better than a histogram bound we compute exactly. A code that gives each input its own subspace escapes the theorem, which locates the problem in the encoder rather than in addition. Tokaformer, the architecture we study, keeps each position's inputs as separate typed tokens through a linear-recurrent routing stage and only then discards the extra tokens. On position-, content-, and track-addressed synthetic tasks, factored input succeeded where summed input failed, and the failure was not explained by sequence length or shared tuning. On aggregation tasks such as counting, neither fusion held a robust advantage, which suggests pooling only the inputs an aggregate can absorb; this hybrid reached similar observed composite accuracy to full factoring (about 0.996) with 35% fewer routing tokens. On WeatherBench 2 and BigEarthNet the ranking of the two fusions depended on the input encoder: with one encoder shared across channels factored fusion won clearly, while with a separate encoder per channel, which binds identity to value before the sum, summed fusion was as good or better, consistent with the proposed mechanism. We summarize the results as bind before you compress: keep inputs separate until the network has selected among them, or until the encoding itself has recorded which is which, and pool wherever selection is unnecessary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.