attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 04

Sinusoidal positions

Instead of learning a table, compute each position's vector from sine and cosine waves of many speeds.

The formula, then the picture

PE(pos, 2i) = sin(pos · ωi)   PE(pos, 2i+1) = cos(pos · ωi)   ωi = 1 / 100002i/d

Each pair of dimensions is one wave. Early pairs wiggle fast, later pairs crawl.

Read a position off the waves

The famous heatmap

Clock hands at different speeds

Why waves and not just numbers

Wavelengths form a geometric ladder

The dot product only cares about distance

PE(p) · PE(p+k) = Σi cos(k · ωi)   (no p anywhere)

A shift is a rotation

“Works at any length” is only half true

Exceptions and sharp edges

Position and meaning share lanes

the model learns to separate them. RoPE avoids the mixing entirely.

Slow lanes are nearly constant

fine: they leave room for content.

Interleaved or split halves?

match the layout used in training.

Precision at huge positions

compute angles in fp32, then cast.

The base 10000 is a choice

long-context models often raise the base.