The real bottleneck
The GPU memory hierarchy
Standard attention's round trips
Tiles and online softmax
One query block, sweeping key blocks
Online softmax, with real numbers
m′ = max(m, max s) ℓ′ = em−m′ℓ + Σ es−m′ O′ = em−m′O + Σ es−m′v out = O / ℓ
What it buys
HBM traffic against sequence length
Backward: recompute instead of store
Causal masking at block level
Three versions
Exceptions and sharp edges
Exact, not approximate
expect tiny numeric differences, not quality loss.fp16 and bf16 only
run attention under bf16 autocast.Head dimension limits
keep head_dim at 64, 128 or 256.Nondeterministic backward
use the deterministic flag when you need bitwise reproducibility.Packed sequences
varlen kernels take cumulative sequence lengths instead of padding.