attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

octlm · day 1 · transformer foundations

Attention, drawn one stroke at a time

Thirteen short chapters that follow one sentence from raw bytes to a predicted next token.

The whole trip, one sentence long

Our running example

How to read the drawings

Shapes along the way

Chapters

chapter 01 · act 1Byte-level BPEText becomes integer ids
chapter 02 · act 1Token embeddingsIds become vectors
chapter 03 · act 2Positional informationOrder is invisible until you add it
chapter 04 · act 2Sinusoidal positionsA ruler made of waves
chapter 05 · act 2RoPERotate the query, rotate the key
chapter 06 · act 3Q, K and VAsk, advertise, hand over
chapter 07 · act 3Scaled dot-productScore, scale, softmax, mix
chapter 08 · act 3Causal maskingNo peeking at the answer
chapter 09 · act 3Multi-head attentionMany small lookups at once
chapter 10 · act 3The n × n matrixWhere the quadratic cost lives
chapter 11 · act 3KV cacheKeep the keys, skip the redo
chapter 12 · act 4The decoder blockAttention and MLP on a residual highway
chapter 13 · act 4Optimization and reproducibilityTrain it, then train it again identically