attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 12

The decoder block

Attention lets tokens talk to each other.

Inside one block

One trip through a pre-norm block

h = x + Attn(Norm(x))     out = h + MLP(Norm(h))

The residual stream is a shared notebook

Pre-norm vs post-norm

The pieces

LayerNorm and RMSNorm

The MLP: widen, bend, narrow

Attention mixes across tokens, the MLP within a token

The whole model

From the last block to the next word

Where the parameters live

Exceptions and sharp edges

Forgetting the final norm

norm → unembed, always.

The stream grows with depth

scale residual-branch init by 1/√(2N), as GPT-2 does.

SwiGLU has three matrices

read hidden size from the config, never assume 4d.

Norm epsilon

compute norms in fp32, keep ε around 1e-5 or 1e-6.

Dropout placement

most large LLMs use no dropout in pretraining at all.