Certified Nonlinear Computing for Scaled FP8 Datapaths
Abstract
Large Language Models (LLMs) demand substantial computation for inference, motivating the adoption of efficient FP8 matrix engines. However, nonlinear operations such as SwiGLU, RMSNorm, and softmax-based attention still commonly rely on wider intermediate arithmetic before FP8 quantization. Premature rounding of these intermediates can alter the final output codes after downstream composition and shared-scale selection. In this paper, we propose a sufficient-state framework for efficient nonlinear computation in scaled-FP8 pipelines. The framework constructs intermediate states that preserve the information needed to determine final element codes and shared scale bytes. We instantiate this principle through directed states for SwiGLU, exact boundary predicates over whole-row statistics for RMSNorm, and interval aggregates with adaptive refinement for softmax attention. These constructions resolve final quantization decisions through consumer composition and shared reductions. For SwiGLU, mathematical verification and exhaustive finite-state replay establish exact outputs under the declared numerical contract; for attention, interval bounds certify output cells when they are uniquely determined. Based on this framework, we design and implement complete hardware endpoints for SwiGLU, RMSNorm, and softmax attention. Hardware experiments in Nangate45 demonstrate that the SwiGLU semantic ROM reduces area by 54–61% and gate energy by 41–53% relative to polynomial controls. Under matched complete-row configurations, RMSNorm reduces area by 4.1% and gate energy by 23.0%, while attention achieves reductions of 58.8% and 79.6%, respectively, relative to their corresponding controls. Joint replacement experiments across three language models on GSM8K, HumanEval, and GPQA-Diamond demonstrate task performance comparable to native execution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.