Removing Redundant Computation in a BDH-GPU-Derived Language Model: Exact Recurrent Training and Conditional Neuron Selection
Abstract
Dragon Hatchling (BDH-GPU) combines strict-past Q=K linear attention with a very wide positive neuronal representation. We study Arm-A, a 17.4M-parameter BDH-GPU-derived language model, and ask which apparent redundancy is actually removable. For attention, an exact packed segmented scan preserves the dense operator while avoiding cross-block score materialization. A predeclared Black- well search improves our original dense path by 1.42×, yet recurrence remains 1.52× faster at T =2048. At a fixed 131,072 positions/update, the recurrent/dense advantage grows from 1.04× at T =512 to 2.12× at T =4096. Applying the same mature executor to canonical 25M BDH still leaves recurrence below dense parity (0.964×), and a 96-update matched BF16 trajectory differs by only 7.05 × 10−4 held-out NLL despite non-bitwise parameter drift. For neuronal sparsity, causal controls show that static-support damage is directional rather than magnitude-only: preserving dense direction at the static-support norm reduces ∆NLL from 0.0906 to 0.00819 relative to the converse norm-matched intervention, with the same ordering in KL. Exact algebraic redundancy is therefore removable now; sparse neuronal computation requires conditional support selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.