modern arch

  1. 00The mapFrom the 2017 block to a 2024 LLM block
  2. 01RoPE, properlyFrequencies, caches, conventions and long context
  3. 02RMSNormRescale, skip the re-centering
  4. 03SwiGLUA gated MLP with a smooth switch
  5. 04Grouped-query attentionMany queries, few keys and values
  6. 05SDPAOne call, three backends
  7. 06FlashAttentionExact attention that respects memory
  8. 07The modern decoder blockEvery upgrade, assembled

chapter 06

FlashAttention

Attention is slow because of memory traffic, not arithmetic.

The real bottleneck

The GPU memory hierarchy

Standard attention's round trips

Tiles and online softmax

One query block, sweeping key blocks

Online softmax, with real numbers

m′ = max(m, max s)   ℓ′ = em−m′ℓ + Σ es−m′   O′ = em−m′O + Σ es−m′v   out = O / ℓ

What it buys

HBM traffic against sequence length

Backward: recompute instead of store

Causal masking at block level

Three versions

Exceptions and sharp edges

Exact, not approximate

expect tiny numeric differences, not quality loss.

fp16 and bf16 only

run attention under bf16 autocast.

Head dimension limits

keep head_dim at 64, 128 or 256.

Nondeterministic backward

use the deterministic flag when you need bitwise reproducibility.

Packed sequences

varlen kernels take cumulative sequence lengths instead of padding.