SMIC: Shared Multiscale Input-Reference Conditioning for Natural Language Inference
Abstract
Natural Language Inference models can be sensitive to localized lexical changes. We introduce Shared Multiscale Input-Reference Conditioning (SMIC), which constructs a single input-dependent multiscale reference and reuses it across Transformer layers. Each layer applies its own multiscale transformation to the shared reference and injects the resulting conditioning feature before the attention update. We instantiate SMIC in a token–context Transformer (TCT) and conduct from-scratch comparisons against decoder-based and MLP baselines across all three ANLI rounds, with additional encoder comparisons on ANLI-R1. TCT-SMIC attains the highest observed ten-seed mean accuracy among the evaluated configurations. The ablations further support the contribution of SMIC. We also evaluate a real-valued Llama-style plug-in and extend deletion/insertion diagnostics to pretrained ModernBERT encoder and Llama decoder models with and without SMIC. Across the evaluated perturbation protocols, TCT-SMIC exhibits comparatively low pooled-attention redistribution and smaller signed accuracy changes. We further establish a finite-sample bound on average subset-mass redistribution in terms of the mean Jensen–Shannon divergence between the restricted and renormalized attention distributions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.