attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 06

Q, K and V

Every token plays three roles at once.

The intuition

A library, very quickly

Making Q, K and V

Same x, three lenses

Q = x WQ   K = x WK   V = x WV     [T × d] · [d × dk] = [T × dk]

Every token gets all three

Matching and mixing

“it” looks for its antecedent

Mixing the values

outi = Σj wij · vj

Why keys and values are separate

Shapes cheat sheet

Self-attention vs cross-attention

Exceptions and sharp edges

Scores are not symmetric

read attention maps row by row: one row is one query.

Tokens attend to themselves

nothing to fix, the residual connection relies on it too.

dk must match

assert q.shape[-1] == k.shape[-1].

Weights are not explanations

treat attention maps as hints, not proof.

Bias or no bias

check the config before loading weights.