WHEN IS A SECOND REPRESENTATION WORTH ACTIVATING? INFORMATION WAVE–PARTICLE DUALITY FOR SEQUENTIAL LEARNING
Abstract
Modern sequence models increasingly combine spectral, temporal, local, and global views, implicitly treating more representation as better. We show why that assumption can fail under finite model capacity and ask a more basic question: when is a second representation worth activating? We introduce Information Wave–Particle Duality (IWPD), where a distributed spectral Wave view and a localized Packet view are alternative lossy representations of the same observed history. Under Bayes log loss, the context-specific value of adding one view after the other is exactly a directional conditional mutual information. For restricted predictors, we derive a finite-capacity decomposition showing that positive population information can still be offset by approximation burden; adding representation cost yields an explicit activation law. We instantiate the Wave branch with a constrained maximum-entropy spectral prior and deploy the theory through history-only prediction of the task risk of Base, Wave, Packet, and Dual actions. Experiments across five datasets reject universal Dual superiority. Most importantly, under an identical-input, identical-split four-action trajectory protocol, the history-only selector attains lower five-seed mean ADE than the condition-wise best fixed action in all 12 highD/NGSIM horizon conditions. Failures outside this setting are retained explicitly. The resulting principle is not “fuse everything,” but activate a representation only when its conditional value survives model and cost constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.