attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 10

The n × n matrix

Every token scores every token, so the score matrix has T² entries per head per layer.

Seeing the square

Add tokens, watch the grid explode

Doubling is quadrupling

Compute and memory

When does attention dominate?

Memory to store the full score matrix

How real systems cope

FlashAttention: never build the whole square

The online softmax trick behind it

Or compute fewer scores

Exceptions and sharp edges

Short prompts: blame the MLP

profile before optimising.

FlashAttention is exact

use it by default, it is not an approximation.

Padding wastes squares

sort by length, pack sequences, or use variable-length kernels.

Decode is linear per token

see chapter 11.

Linear attention trades quality

hybrids mix a few full-attention layers back in.