B-Net: Learning Where to Compress in Byte-Level Language Models
Abstract
Hierarchical byte-level Transformers operate directly on bytes and group them into patches for efficient computation. Their patcher must identify useful boundaries without language-specific rules while keeping the resulting computational budget predictable. Space-based patchers encode language-specific structure, while learned patchers controlled by a soft ratio loss can deviate substantially from requested patch lengths. We introduce B-Net, a hierarchical byte-level Transformer that separates learned boundary placement from explicit compression control. A causal convolutional scorer learns input-dependent boundaries without a tokenizer or manually specified language features, while an adaptive threshold controller realizes the requested average patch rate. On English, the learned patcher recovers the quality of strong space-based segmentation without receiving spaces as boundary rules. The same patching mechanism remains effective in separately trained Chinese models. We compare B-Net with state-of-the-art byte-level Transformers under a FLOP- and data-matched protocol. B-Net outperforms the strongest byte-level baseline in each setting by 3.4 percentage points in mean downstream accuracy on English and 4.1 points in XWinograd accuracy on Chinese. It improves English WikiText performance from 0.85 to 0.81 bits per byte and matches the best Chinese validation bits per byte. These results show that learned byte patching can combine reliable rate control with superior predictive performance without manually specified language-boundary rules.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.