attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 11

KV cache

Generation adds one token at a time.

Without and with a cache

No cache: redo everything, every step

With a cache: compute one row, append, look up the rest

Work saved

Why K and V, but not Q

Old queries are never asked again

The cost: memory

KV cache size calculator

bytes = 2 × layers × kv_heads × head_dim × tokens × batch × bytes_per_value

Sharing K and V shrinks the cache

Prefill, then decode

Exceptions and sharp edges

Editing an earlier token

truncate the cache at the edit and recompute from there.

Position offset

pass cache length as the position offset.

Ragged batches

paged attention stores the cache in fixed-size blocks.

Evicting old entries

keep the first few tokens plus a recent window.

Quantised caches

measure on long-context tasks before shipping.

Shared prefixes

prompt caching reuses the prefix's KV blocks.