attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 02

Token embeddings

An id is a name tag, not a quantity.

The lookup

Seven ids fetch seven rows

Why it counts as a matrix multiply

A map of meaning

Neighbours in embedding space

Vectors as colour strips

Cosine similarity between them

Cosine similarity is an angle

cos θ = a · b ÷ (‖a‖ ‖b‖)

Where the map comes from

Training pulls the points into clusters

The table is big, so some models share it

Exceptions and sharp edges

One vector per token, any context

attention mixes in context later. That is the point of the rest of the model.

Ids are not quantities

always look up a vector, never use the id as a value.

Rare tokens barely move

sensible vocab size, and tokenizer data that matches model data.

Out-of-range id

check that the tokenizer and model agree on V, including special tokens.

Padding has a row too

mask pad positions in attention and in the loss.

Scale matters

follow your architecture's convention, then keep it consistent.