modern arch

  1. 00The mapFrom the 2017 block to a 2024 LLM block
  2. 01RoPE, properlyFrequencies, caches, conventions and long context
  3. 02RMSNormRescale, skip the re-centering
  4. 03SwiGLUA gated MLP with a smooth switch
  5. 04Grouped-query attentionMany queries, few keys and values
  6. 05SDPAOne call, three backends
  7. 06FlashAttentionExact attention that respects memory
  8. 07The modern decoder blockEvery upgrade, assembled

chapter 04

Grouped-query attention

Every query head in Day 1 had its own keys and values.

The spectrum

MHA, GQA, MQA

Making the shapes line up

[B, hkv, T, dh] → repeat_interleave(h / hkv, dim=1) → [B, h, T, dh]

Which KV head does query head i use?

What it saves

KV cache size against context length

bytes = 2 · layers · hkv · dh · tokens · 2

Decode is a memory-reading contest

The trade-off

Converting an MHA checkpoint

The code

Exceptions and sharp edges

h must divide by hkv

choose h_kv from the divisors of h.

repeat vs repeat_interleave

use repeat_interleave or the expand-reshape trick.

More GPUs than KV heads

replicate KV heads across GPUs.

Caching the repeated tensor

cache h_kv heads, repeat on the fly or let the kernel handle it.

Kernels can skip the copy

pass h_kv-shaped K and V when the kernel supports it.