Think Neuron: Belief Revision as a Sequence Transition
Abstract
Most sequence architectures transform point representations. Probabilistic extensions usually attach uncertainty to parameters, latent variables, layer outputs, embeddings, or attention operators. We instead place the probabilistic object at the neural unit and make local belief revision the sequence transition itself. Elemental Variational Expanse (EVE), the variational distributional neuron used here, defines an input-conditioned local posterior. Affine-EVE turns prior–posterior revision into an analytical vector-valued transition. For diagonal Gaussian prior and evidence, the transition is exactly . These affine transitions compose associatively; we call this property Variational Chain Closure (VCC). EVER is the resulting sequential architecture. Across language modeling, semantic classification, audio, biological sequences, and structured composition, EVER achieves the lowest recorded headline value on seven of eight generalist tasks. We then use two WikiText-2 controls to isolate the source of the gain. In a frozen 6.6514M-parameter system benchmark, analytical prior–evidence precision fusion outperforms both a parameter-matched Transformer and a causal cross-attention model on all five paired seeds in NLL. In a stricter control with exactly matched architecture and parameter count, the precision-fusion model retains lower NLL: versus for a direct deterministic transition (5/5 paired seeds; paired -test ). Brier score, accuracy, and top-5 accuracy also favor precision fusion on all five pairs. A frozen FineWeb scaling suite from 10M to 120M parameters further shows EVER outperforming the matched Transformer at every tested size, with Mamba-2 as a strong sequential reference. Same-checkpoint interventions show direct use of prior–posterior quantities, while retrained ablations show that optimization can also reach effective alternative solutions. Together, these results support local belief revision as an explicit, reusable inductive bias for sequence computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.