acceptodds
Under review as a conference paper at ICLR 2027

WHEN IS A SECOND REPRESENTATION WORTH ACTIVATING? INFORMATION WAVE–PARTICLE DUALITY FOR SEQUENTIAL LEARNING

Abstract

Modern sequence models increasingly combine spectral, temporal, local, and global views, implicitly treating more representation as better. We show why that assumption can fail under finite model capacity and ask a more basic question: when is a second representation worth activating? We introduce Information Wave–Particle Duality (IWPD), where a distributed spectral Wave view and a localized Packet view are alternative lossy representations of the same observed history. Under Bayes log loss, the context-specific value of adding one view after the other is exactly a directional conditional mutual information. For restricted predictors, we derive a finite-capacity decomposition showing that positive population information can still be offset by approximation burden; adding representation cost yields an explicit activation law. We instantiate the Wave branch with a constrained maximum-entropy spectral prior and deploy the theory through history-only prediction of the task risk of Base, Wave, Packet, and Dual actions. Experiments across five datasets reject universal Dual superiority. Most importantly, under an identical-input, identical-split four-action trajectory protocol, the history-only selector attains lower five-seed mean ADE than the condition-wise best fixed action in all 12 highD/NGSIM horizon conditions. Failures outside this setting are retained explicitly. The resulting principle is not “fuse everything,” but activate a representation only when its conditional value survives model and cost constraints.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.