The spectrum
MHA, GQA, MQA
Making the shapes line up
[B, hkv, T, dh] → repeat_interleave(h / hkv, dim=1) → [B, h, T, dh]
Which KV head does query head i use?
What it saves
KV cache size against context length
bytes = 2 · layers · hkv · dh · tokens · 2
Decode is a memory-reading contest
The trade-off
Converting an MHA checkpoint
The code
Exceptions and sharp edges
h must divide by hkv
choose h_kv from the divisors of h.repeat vs repeat_interleave
use repeat_interleave or the expand-reshape trick.More GPUs than KV heads
replicate KV heads across GPUs.Caching the repeated tensor
cache h_kv heads, repeat on the fly or let the kernel handle it.Kernels can skip the copy
pass h_kv-shaped K and V when the kernel supports it.