FlashBoB-2: Algorithmic and Hardware Optimizations for Attention JVPs and Double Backward
Abstract
Jacobian–vector products (JVPs) and second-order derivatives of attention arise in training consistency models and learning through gradient updates. FlashAttention primarily optimizes attention evaluation and first-order reverse-mode differentiation, while specialized kernels such as FlashBoB and jvp_flash_attention support forward-mode or second-order operations. Building on FlashAttention-3/-4, we introduce FlashBoB-2 for attention JVPs and backward-over-backward (BoB) on Hopper and Blackwell GPUs, combining a reformulated BoB computation with schedules designed for each GPU. First, we optimize FlashBoB's computational graph, reducing the matrix-multiply count for exact BoB from eighteen to fifteen without storing intermediates and with on-chip memory independent of sequence length . Second, on Hopper, separate warpgroups build intermediate results and accumulate outputs, so each output product starts as soon as its inputs are ready. Third, on Blackwell, we reuse Tensor Memory once each value has been read for the last time and keep intermediate results in shared memory. In causal BF16 benchmarks with head dimension 128, FlashBoB-2 achieves BoB speedups of – on H100 and – on B200 over FlashBoB. For non-causal JVPs at 1K–128K against two baselines, speedups are – on H100 and – on B200.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.