RONDO: Interleaved Datapaths for Division-Bearing Neural Recurrences
Abstract
Regulated domains such as finance and healthcare, where law requires automated decisions affecting individuals to be explainable, need machine learning models with auditable predictions. This requirement motivates interpretable architectures such as continued-fraction networks (CoFrNets) and Chebyshev Kolmogorov-Arnold networks. Each architecture’s functional form exposes the reasoning behind every prediction. Internally, both architectures evaluate a second-order recurrence where each step depends on the two before it. However, this recurrence creates a bottleneck that limits existing hardware (GPUs and vector processors) to as little as 29% of its matrix-multiply rate and 44.5% of ideal parallel speedup. The bottleneck stems from serial dependence, as each step must wait for the previous result. On a five-stage pipeline, fast floating-point hardware therefore idles four fifths of the time, reducing throughput fivefold. This loss limits how quickly these regulated domains can deliver auditable decisions. To address this bottleneck, we propose RONDO, a hardware datapath that evaluates second-order recurrences with input-dependent coefficients, enabling interpretable models to deliver auditable decisions at low latency, high throughput, and low power and area. RONDO computes each recurrence with multiply-adds and at most one division. To keep that divider busy despite the recurrence’s serial dependence, RONDO time-multiplexes independent evaluations onto it, advancing each in turn. We evaluate RONDO on a Kintex-7 FPGA against existing hardware, including an RTX 5090 GPU, the Ara vector processor, and PolyKAN accelerator that divides at every recurrence step. When scaled to a common process node using published scaling factors, RONDO offers 43× lower latency and 9.5× lower energy per evaluation than the best-performing baseline (RTX 5090 GPU). RONDO also achieves 102 × and 105× the throughput of Ara and PolyKAN accelerator, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.