attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 07

Scaled dot-product attention

One line of maths runs every attention head in every modern LLM.

Attention(Q, K, V) = softmax( Q Kᵀ / √dk + mask ) · V

The whole computation

Six steps, one matrix at a time

Every row is chapter 6 again

Why divide by √dk

Dot products spread out as width grows

Large scores freeze softmax

Softmax, carefully

Softmax step by step, with the overflow trap

softmax(s)i = esi − max(s) ÷ Σj esj − max(s)   (same answer, no overflow)

What a dot product measures

The output is a blend, always inside the values

Exceptions and sharp edges

A fully masked row

make sure every query can see at least one key, usually itself.

−1e9 in fp16

use the dtype's minimum value instead of a magic number.

All scores equal

normal at initialisation; training breaks the ties.

Attention dropout

turn dropout off for evaluation with model.eval().

Not always 1/√dk

do not change the scale of a pretrained model.