attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 13

Optimization and reproducibility

Training repeats one loop millions of times: predict, measure the miss, compute gradients, nudge every weight.

The loop

One training step, six moves

Cross-entropy: the price of surprise

loss = − log p(target)   ·   perplexity = emean loss

Stepping downhill

Learning rate: too small, just right, too big

SGD vs Adam in a narrow valley

Warmup, then cosine decay

Gradient clipping

g ← g · min(1, c / ‖g‖)

Reading loss curves

Reproducibility

Same seed, different seed, nondeterministic kernels

The reproducibility checklist

Exceptions and sharp edges

Atomic adds on GPU

torch.use_deterministic_algorithms(True), at some speed cost.

Different hardware

record versions; expect “statistically same”, not “bitwise same”, across machines.

Data loader workers

seed each worker from the base seed and worker id.

Resuming loses state

checkpoint optimizer, scheduler, RNG states and the data position.

Changing batch size

retune LR, or scale it with batch size and re-check.

Loss spikes

clip gradients, lower peak LR, or rewind and skip the batch.