attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 09

Multi-head attention

One head produces one set of weights per token.

Single head vs multi-head

One softmax, two needs

Split the width, do not multiply it

What different heads learn

Four heads, four habits

The same four heads as matrices

Putting heads back together

Concatenate, then mix with WO

MultiHead(x) = concat(head₁, …, headh) · WO

More heads, same parameter bill

Heads can share keys and values

Exceptions and sharp edges

dmodel must divide by h

pick h so d_model / h is a whole number, usually 64 or 128.

Heads that are too narrow

keep d_head around 64 to 128 and scale heads with width.

Many heads are redundant

useful for compression, but you still need them during training.

Attention sinks

keep the first tokens in the cache when trimming long contexts.

The reshape bug

check one head's output against a single-head reference.